Acoustic device and method for determining its transfer function
The acoustic device enhances hearing experience in open-type earphones by estimating and controlling noise reduction signals using transfer functions, ensuring accurate noise cancellation at the target spatial position.
Patent Information
- Application Number
- JP2024502215
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-19
- Filing Date
- 2022-03-03
- Publication Date
- 2025-08-06
- Estimated Expiration
- 2042-03-03
AI Technical Summary
In open-type earphones, the feedback microphone and the target spatial position (e.g., the eardrum) are not in a pressure field environment, leading to inaccurate estimation of sound wave signals for active noise reduction, which degrades the user's hearing experience.
An acoustic device with a sound generation unit, a first detector, and a processor that estimates a second residual signal at a target spatial position using transfer functions and noise reduction control signals, while being fixed near the ear without blocking the ear canal, allowing for accurate noise reduction.
The device effectively reduces environmental noise at the target spatial position, improving the user's hearing experience by accurately controlling noise reduction signals.
Smart Images

Figure 0007719546000016 
Figure 0007719546000017 
Figure 0007719546000018
Abstract
Description
[Technical Field]
[0001] The present application relates to the technical field of acoustics, and in particular to an acoustic device and a method for determining its transfer function.
[0002] [Incorporated by reference] This application claims priority to a Chinese patent application with application number 202111408329.8 filed on November 19, 2021, the entire contents of which are incorporated herein by reference. [Background technology]
[0003] In conventional earphones, during operation, the feedback microphone used for active noise reduction and the target spatial position (e.g., the eardrum of a human ear) are located in a pressure field, and the sound pressure distribution at each position in the sound field is considered to be uniform, so the signal collected by the feedback microphone can directly reflect the sound heard by the human ear. However, in the case of open-type earphones, the environment in which the feedback microphone and the target spatial position (e.g., the eardrum of a human ear) are located is not a pressure field environment, so the signal received by the feedback microphone cannot directly reflect the signal at the target spatial position (e.g., the eardrum of a human ear). Furthermore, the backward sound wave signal emitted from the speaker for active noise reduction cannot be accurately estimated, which reduces the effectiveness of active noise reduction and further degrades the user's hearing experience. Summary of the Invention [Problem to be solved by the invention]
[0004] Therefore, it is desirable to provide an acoustic device that can open both ears of a user to improve the user's hearing experience. [Means for solving the problem]
[0005] An embodiment of the present disclosure provides an acoustic device including a sound generation unit, a first detector, a processor, and a fixed structure, wherein the sound generation unit generates a first audio signal based on a noise reduction control signal, the first detector picks up a first residual signal including a residual noise signal in which environmental noise and the first audio signal are superimposed at the first detector, the processor estimates a second residual signal at a target spatial position based on the first audio signal and the first residual signal, and updates the noise reduction control signal based on the second residual signal, and the fixed structure fixes the acoustic device at a position near the user's ear but not blocking the user's ear canal, such that the target spatial position is closer to the user's ear canal than the first detector.
[0006] In some embodiments, estimating a second residual signal at a target spatial position based on the first audio signal and the first residual signal includes obtaining a first transfer function between the sound production unit and the first detector, a second transfer function between the sound production unit and the target spatial position, a third transfer function between an environmental noise source and the first detector, and a fourth transfer function between the environmental noise source and the target spatial position, and estimating the second residual signal at the target spatial position based on the first transfer function, the second transfer function, the third transfer function, the fourth transfer function, the first audio signal, and the first residual signal.
[0007] In some embodiments, obtaining a first transfer function between the sound producing unit and the first detector, a second transfer function between the sound producing unit and the target spatial position, a third transfer function between an environmental noise source and the first detector, and a fourth transfer function between the environmental noise source and the target spatial position includes obtaining the first transfer function, and determining the second transfer function, the third transfer function, and the fourth transfer function based on the first transfer function and each mapping relationship between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function.
[0008] In some embodiments, each of the mapping relationships between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function is generated based on test data in different wearing scenes of the acoustic device.
[0009] In some embodiments, obtaining a first transfer function between the sound production unit and the first detector, a second transfer function between the sound production unit and the target spatial location, a third transfer function between an environmental noise source and the first detector, and a fourth transfer function between the environmental noise source and the target spatial location includes obtaining the first transfer function, inputting the first transfer function into a trained neural network, and obtaining an output of the trained neural network as the second transfer function, the third transfer function, and the fourth transfer function.
[0010] In some embodiments, obtaining the first transfer function includes calculating the first transfer function based on the noise reduction control signal and the first residual signal.
[0011] In some embodiments, the acoustic device further includes a distance sensor that detects a distance from the acoustic device to the user's ear, and the processor further determines the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the distance.
[0012] In some embodiments, estimating a second residual signal at a target spatial position based on the first audio signal and the first residual signal includes obtaining a first transfer function between the sound production unit and the first detector, a second transfer function between the sound production unit and the target spatial position, and a fifth transfer function reflecting a relationship between an environmental noise source, the first detector, and the target spatial position; and estimating the second residual signal at the target spatial position based on the first transfer function, the second transfer function, the fifth transfer function, the first audio signal, and the first residual signal.
[0013] In some embodiments, the first transfer function and the second transfer function have a first mapping relationship, and the fifth transfer function and the first transfer function have a second mapping relationship.
[0014] In some embodiments, estimating a second residual signal at a target spatial location based on the first audio signal and the first residual signal includes obtaining a first transfer function between the sound production unit and the first detector, and estimating the second residual signal at the target spatial location based on the first transfer function, the first audio signal, and the first residual signal.
[0015] In some embodiments, the target spatial location is the user's eardrum location.
[0016] An embodiment of the present disclosure also provides a method for determining a transfer function of an acoustic device, the acoustic device including a sound generation unit, a first detector, a processor, and a fixed structure, the fixed structure fixing the acoustic device to a position near an ear of a subject but not blocking the ear canal of the subject, the method including the steps of: acquiring, in a scene without environmental noise, a first signal emitted by the sound generation unit based on a noise reduction control signal; and a second signal picked up by the first detector, the second signal including a residual noise signal transmitted to the first detector by the first signal; determining a first transfer function between the sound generation unit and the first detector based on the first signal and the second signal; and detecting a transfer function of the sound generation unit using a second detector installed at a target spatial position closer to the ear canal of the subject than the first detector. The present invention provides a method including the steps of: acquiring a third signal picked up, the third signal including a residual noise signal transmitted to the target spatial position by the first signal; determining a second transfer function between the sound producing unit and the target spatial position based on the first signal and the third signal; acquiring a fourth signal picked up by the first detector and a fifth signal picked up by the second detector in a scene where the environmental noise is present and the sound producing unit does not transmit any signal; determining a third transfer function between an environmental noise source and the first detector based on the environmental noise and the fourth signal; and determining a fourth transfer function between the environmental noise source and the target spatial position based on the environmental noise and the fifth signal.
[0017] In some embodiments, the method further includes determining multiple sets of transfer functions for different wearing scenes or different subjects, each set of transfer functions including a corresponding first transfer function, a second transfer function, a third transfer function, and a fourth transfer function; and determining a mapping relationship between the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the multiple sets of transfer functions.
[0018] In some embodiments, determining a mapping relationship among the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the plurality of sets of transfer functions includes: training a neural network using the plurality of sets of transfer functions as training samples; and determining the mapping relationship among the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function from the trained neural network.
[0019] In some embodiments, the mapping relationship between the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function includes a first mapping relationship between the first transfer function and the second transfer function, a ratio between the third transfer function and the fourth transfer function, and a second mapping relationship between the third transfer function and the fourth transfer function and the first transfer function.
[0020] In some embodiments, the first transfer function is positively correlated with a ratio of the second signal to the first signal, the second transfer function is positively correlated with a ratio of the third signal to the first signal, the third transfer function is positively correlated with a ratio of the fourth signal to the environmental noise, and the fourth transfer function is positively correlated with a ratio of the fifth signal to the environmental noise.
[0021] In some embodiments, determining a mapping relationship among the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the plurality of sets of transfer functions includes obtaining a distance from the acoustic device to a corresponding ear of the subject for the different wearing scenes or the different subjects, and determining a mapping relationship among the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the distance and the plurality of sets of transfer functions.
[0022] In some embodiments, the target spatial location is the location of the subject's eardrum.
[0023] Additional features of the present disclosure can be set forth in part in the description that follows. Additional features of the present disclosure will become apparent to those skilled in the art upon study of the following description and corresponding drawings, or upon understanding the manufacture or operation of the embodiments. Features of the present disclosure can be realized or attained by practicing or using various aspects of the methods, tools, and combinations that are described in the following detailed examples.
[0024] The present disclosure will be further illustrated by exemplary embodiments, which will be described in detail with reference to the drawings, which are not limiting, and in which like reference numerals represent like structures. [Brief explanation of the drawings]
[0025] [Figure 1] FIG. 1 is a schematic block diagram of an exemplary acoustic device according to some embodiments of the present disclosure. [Figure 2] 1 is a schematic diagram of an acoustic device in a worn state according to some embodiments of the present disclosure. FIG. [Figure 3] 1 is a flowchart of an exemplary method for reducing noise in an acoustic device according to some embodiments of the present disclosure. [Figure 4] 1 is an exemplary flowchart of a method for determining a transfer function of an acoustic device according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0026] In order to more clearly describe the technical means of the embodiments of the present disclosure, the drawings necessary for describing the embodiments will be briefly described below. Obviously, the drawings described below are merely some examples or embodiments of the present disclosure, and those skilled in the art can apply the present disclosure to other similar scenarios based on these drawings without any creative effort. It should be understood that these exemplary embodiments are merely intended to enable those skilled in the art to better understand and implement the present disclosure, and do not limit the scope of the present disclosure in any way. Unless otherwise clear from the linguistic environment or otherwise described, the same numbers in the drawings indicate the same structures or operations.
[0027] It should be understood that the terms "system," "device," "unit," and / or "module" used in this disclosure are meant to distinguish between different levels of components, elements, parts, sections, or assemblies. However, other terms may be substituted for the terms provided they achieve the same purpose.
[0028] As used herein and in the claims, unless the context clearly dictates otherwise, the words "a," "one," "one," and / or "the" do not specifically refer to the singular but may include the plural. Generally, the terms "comprise" and "containing" merely indicate the inclusion of the specified steps and elements, and these steps and elements are not an exclusive listing, and a method or apparatus may also include other steps or elements. The term "based on" means "based at least in part on." The term "one embodiment" refers to "at least one embodiment." The term "another embodiment" refers to "at least one other embodiment."
[0029] In the description of this disclosure, the terms "first," "second," "third," "fourth," etc. are used for descriptive purposes only and should not be understood as indicating or suggesting relative importance or implicitly designating the quantity of the indicated technical features. Thus, a feature qualified by "first," "second," "third," or "fourth" can explicitly or implicitly indicate the inclusion of at least one of that feature. In the description of this disclosure, unless otherwise clearly and specifically limited, "plurality" means at least two, e.g., two, three, etc.
[0030] In the present disclosure, unless otherwise clearly specified or limited, the terms "connected," "fixed," etc. should be understood in a broad sense. For example, unless otherwise clearly limited, the term "connected" may mean a fixed connection, a detachable connection, or an integral connection, a mechanical connection, or an electrical connection, a direct connection, an indirect connection via an intermediate medium, internal communication between two elements, or interaction between two elements. Those skilled in the art can understand the specific meanings of the above terms in the present disclosure depending on the specific circumstances.
[0031] This disclosure uses flowcharts to describe operations performed by systems according to embodiments of the present disclosure. It should be understood that preceding or following operations are not necessarily performed in exact order. Instead, steps may be processed in reverse order or simultaneously. Additionally, other operations may be added to these processes, or one or more steps may be removed from these processes.
[0032] An open-type acoustic device (e.g., open-type acoustic earphones) is an acoustic device that allows a user's ears to be open. An open-type acoustic device can fix a speaker near the user's ear but in a position that does not block the user's ear canal using a fixing structure (e.g., ear hooks, head hooks, temples, etc.). When a user uses an open-type acoustic device, external environmental noise is also heard by the user, resulting in a poor hearing experience for the user. For example, in a place with loud external environmental noise (e.g., a street or a tourist spot), when a user plays music using an open-type acoustic device, the external environmental noise directly enters the user's ear canal, causing the user to hear the loud environmental noise, which interferes with the user's music listening experience.
[0033] Active noise reduction can improve the auditory experience of users when using an acoustic device. However, in the case of an open-type acoustic device, the environment where the feedback microphone and the target spatial position (e.g., the eardrum or basilar membrane of a human ear) are located is not a pressure field environment, so the signal received by the feedback microphone cannot directly reflect the signal at the target spatial position. Furthermore, the backward sound wave signal emitted from the speaker cannot be accurately feedback-controlled, and the active noise reduction function cannot be successfully implemented.
[0034] To solve the above problems, an embodiment of the present disclosure provides an acoustic device. The acoustic device includes a sound generation unit, a first detector, and a processor. The sound generation unit generates a first audio signal based on a noise reduction control signal. The first detector picks up a first residual signal. The first residual signal may include a residual noise signal obtained by superimposing environmental noise and the first audio signal in the first detector. The processor estimates a second residual signal at a target spatial position based on the first audio signal and the first residual signal, and updates the noise reduction control signal for controlling the sound generation by the sound generation unit based on the second residual signal. A fixing structure fixes the acoustic device in a position near a user's ear but not blocking the user's ear canal, so that the target spatial position is closer to the user's ear canal than the first detector.
[0035] In an embodiment of the present disclosure, the processor can accurately estimate the second residual signal at the target spatial position using the transfer functions among the sound-producing unit, the first detector, the environmental noise source, and the target spatial position and / or the mapping relationship between each transfer function. Furthermore, the processor can accurately control the generation of the noise reduction signal by the sound-producing unit, effectively reduce the environmental noise in the user's ear canal (e.g., the target spatial position), realize active noise reduction of the acoustic device, and improve the user's hearing experience in the process of using the acoustic device.
[0036] Hereinafter, an acoustic device and a method for determining a transfer function thereof according to an embodiment of the present disclosure will be described in detail with reference to the drawings.
[0037] FIG. 1 is a schematic diagram of an exemplary audio device according to some embodiments of the present disclosure. In some embodiments, audio device 100 may be an open-type audio device capable of active noise reduction for external noise. In some embodiments, audio device 100 may be earphones, glasses, an augmented reality (AR) device, a virtual reality (VR) device, or the like. As shown in FIG. 1 , audio device 100 may include a sound output unit 110, a first detector 120, and a processor 130. In some embodiments, sound output unit 110 may generate a first audio signal based on a noise reduction control signal. First detector 120 may pick up a first residual signal in which environmental noise and the first audio signal are superimposed, convert the picked-up first residual signal into an electrical signal, and transmit the electrical signal to processor 130 for processing. Processor 130 may be coupled (e.g., electrically connected) to first detector 120 and audio output unit 110. The processor 130 may receive and process the electrical signal transmitted from the first detector 120. For example, the processor 130 may estimate a second residual signal at a target spatial position based on the first audio signal and the first residual signal, and then update a noise reduction control signal for controlling the production of the production unit 110 based on the second residual signal. The production unit 110 generates an updated noise reduction signal in response to the updated noise reduction control signal, thereby achieving active noise reduction.
[0038] The sound generation unit 110 may be configured to output an audio signal. For example, the sound generation unit 110 may output a first audio signal based on a noise reduction control signal. As another example, the sound generation unit 110 may output a voice signal based on a voice control signal. In some embodiments, an audio signal (e.g., a first audio signal or an updated first audio signal) generated by the sound generation unit 110 based on the noise reduction control signal may be referred to as a noise reduction signal. The noise reduction signal generated by the sound generation unit 110 can reduce or cancel environmental noise transmitted to a target spatial position (e.g., a position in the user's ear canal, such as the eardrum or basilar membrane), thereby realizing active noise reduction in the acoustic device 100 and thereby improving the user's hearing experience during use of the acoustic device 100.
[0039] In the present disclosure, the noise reduction signal may be an audio signal that is out of phase or substantially out of phase with the environmental noise, and the sound waves of the noise reduction signal cancel out some or all of the sound waves of the environmental noise, thereby achieving active noise reduction. As can be understood, a user can select the degree of active noise reduction according to their actual needs. For example, the degree of active noise reduction can be adjusted by adjusting the amplitude of the noise reduction signal. In some embodiments, the absolute value of the phase difference between the phase of the noise reduction signal and the phase of the environmental noise at the target spatial location can be within a preset phase range. The preset phase range can be between 90 and 180 degrees. The absolute value of the phase difference between the phase of the noise reduction signal and the phase of the environmental noise at the target spatial location can be adjusted within the range according to the user's needs. For example, if a user does not want to be disturbed by sounds in the surrounding environment, the absolute value of the phase difference can be a large value, such as 180 degrees, i.e., the phase of the noise reduction signal is opposite to the phase of the environmental noise at the target spatial location. As another example, if a user wants to be sensitive to the surrounding environment, the absolute value of the phase difference can be set to a small value, such as 90 degrees. Note that the greater the amount of ambient sound (i.e., environmental noise) that the user wishes to pick up, the closer the absolute value of the phase difference may be to 90 degrees, and the less ambient sound that the user wishes to pick up, the closer the absolute value of the phase difference may be to 180 degrees. In some embodiments, when the phase of the noise-reduced signal and the phase of the environmental noise at the target spatial position satisfy a certain condition (e.g., the phases are opposite), the amplitude difference between the amplitude of the environmental noise at the target spatial position and the amplitude of the noise-reduced signal may be within a predetermined amplitude range. For example, if the user does not want to be disturbed by ambient sound, the amplitude difference may be a small value such as 0 dB, i.e., the amplitude of the noise-reduced signal is equal to the amplitude of the environmental noise at the target spatial position. As another example, if the user wishes to be sensitive to the ambient sound, the amplitude difference may be a large value, for example, approximately equal to the amplitude of the environmental noise at the target spatial position.In addition, the more sound in the surrounding environment that the user wants to pick up, the closer the amplitude difference is to the amplitude of the environmental noise at the target spatial position, and the less sound in the surrounding environment that the user wants to pick up, the closer the amplitude difference is to 0 dB.
[0040] In some embodiments, when the user wears the acoustic device 100, the sound generation unit 110 may be located near the user's ear. In some embodiments, based on the operating principle of the sound generation unit 110, the sound generation unit 110 may include one or more of an electrodynamic speaker (e.g., a moving coil speaker), a magnetic speaker, an ion speaker, an electrostatic speaker (or a capacitor speaker), a piezoelectric speaker, etc. In some embodiments, depending on the propagation method of the sound output by the sound generation unit 110, the sound generation unit 110 may include an air conduction speaker and / or a bone conduction speaker. In some embodiments, if the sound generation unit 110 is a bone conduction speaker, the target spatial position may be the position of the user's basilar membrane. If the sound generation unit 110 is an air conduction speaker, the target spatial position may be the position of the user's eardrum, so that the acoustic device 100 has a high active noise reduction effect.
[0041] In some embodiments, the number of sound generation units 110 may be one or more. When the number of sound generation units 110 is one, the sound generation unit 110 may output a noise reduction signal to remove environmental noise and transmit audio information (e.g., device media audio, far-end call audio) that the user needs to hear to the user. For example, when the number of sound generation units 110 is one and the sound generation unit 110 is an air conduction speaker, the air conduction speaker may output a noise reduction signal to remove environmental noise. In this case, the noise reduction signal may be sound waves (i.e., air vibrations), which are transmitted to a target spatial location through the air and can cancel out the environmental noise at the target spatial location. At the same time, the air conduction speaker may also transmit audio information that the user needs to hear to the user. As another example, when the number of sound generation units 110 is one and the sound generation unit 110 is a bone conduction speaker, the bone conduction speaker may output a noise reduction signal to remove environmental noise. In such a case, the noise reduction signal may be a vibration signal (e.g., vibration of the speaker housing), which may be transmitted to the user's basilar membrane via the skeleton or tissue and cancel out environmental noise at the user's basilar membrane. At the same time, the bone conduction speaker may also transmit audio information that the user needs to hear to the user. When there are multiple sound production units 110, some of the multiple sound production units 110 may output noise reduction signals to remove environmental noise, and other may transmit audio information that the user needs to hear (e.g., device media audio, far-end call audio) to the user. For example, when there are multiple sound production units 110 and the sound production units 110 include a bone conduction speaker and an air conduction speaker, the air conduction speaker may output sound waves to reduce or remove environmental noise, and the bone conduction speaker may transmit audio information that the user needs to hear to the user. Compared to an air conduction speaker, a bone conduction speaker can transmit mechanical vibrations directly to the user's auditory nerve via the user's body (e.g., skeleton, skin tissue, etc.), resulting in less interference with the air conduction microphone that picks up environmental noise during this process.
[0042] It should be noted that the sound production unit 110 may be an independent functional device or may be part of a single device capable of realizing multiple functions. By way of example only, the sound production unit 110 may be integrated and / or formed integrally with the processor 130. In some embodiments, when there are multiple sound production units 110, the arrangement of the multiple sound production units 110 may include a linear array (e.g., a straight line, a curved line), a planar array (e.g., a regular and / or irregular shape, such as a cross, a mesh, a circle, a ring, a polygon, etc.), a three-dimensional array (e.g., a cylindrical, a spherical, a hemispherical, a polyhedral, etc.), or any combination thereof, and is not limited in the present disclosure. In some embodiments, the sound production unit 110 may be located at the left ear and / or the right ear of the user. For example, the sound production unit 110 may include a first sub-speaker and a second sub-speaker. The first sub-speaker may be located at the user's left ear, and the second sub-speaker may be located at the user's right ear. The first sub-speaker and the second sub-speaker may be activated simultaneously, or may be controlled so that only one of them is activated. In some embodiments, the sound production unit 110 may be a speaker with a directional sound field, the main lobe of which is directed toward the user's ear canal.
[0043] The first detector 120 may be configured to pick up an audio signal. For example, the first detector 120 may pick up a user's voice signal. As another example, the first detector 120 may pick up a first residual signal. In some embodiments, the first residual signal may include a residual noise signal obtained by superimposing environmental noise and the first audio signal (i.e., the noise-reduced signal) generated by the sound production unit 110 in the first detector 120. In other words, the first detector 120 can simultaneously pick up the environmental noise and the noise-reduced signal generated by the sound production unit 110. Furthermore, the first detector 120 can convert the first residual signal into an electrical signal and transmit it to the processor 130 for processing.
[0044] In this disclosure, environmental noise refers to a combination of various external sounds in the environment in which a user is located. By way of example only, environmental noise may include one or more of traffic noise, industrial noise, construction noise, social noise, etc. In some embodiments, traffic noise may include, but is not limited to, automobile driving noise, horn honking noise, etc. Industrial noise may include, but is not limited to, factory power machinery operating noise, etc. Construction noise may include, but is not limited to, power machinery digging noise, drilling noise, stirring noise, etc. Social noise may include, but is not limited to, crowd noise, entertainment advertising noise, crowd noise, household appliance noise, etc.
[0045] In some embodiments, the environmental noise may include the user's voice. For example, the first detector 120 may pick up environmental noise depending on the call state of the acoustic device 100. When the acoustic device 100 is not in a call state, the user's voice may be considered environmental noise, and the first detector 120 may simultaneously pick up the user's voice and other environmental noise. When the acoustic device 100 is in a call state, the user's voice may not be considered environmental noise, and the first detector 120 may pick up environmental noise other than the user's voice. For example, the first detector 120 may pick up noise emitted from a noise source that is at least a certain distance (e.g., 0.5 meters or 1 meter) away from the first detector 120. As another example, the first detector 120 may pick up noise that is significantly different from the user's voice (e.g., the difference in frequency, volume, or sound pressure is greater than a certain threshold).
[0046] In some embodiments, the first detector 120 may be located near the user's ear canal to pick up environmental noise and / or the first audio signal transmitted to the user's ear canal. For example, when the user is wearing the acoustic device 100, the first detector 120 may be located on the side of the sound production unit 110 facing the user's ear canal (as shown by the first detector 220 and the sound production unit 210 in FIG. 2). In some embodiments, the first detector 120 may be located at the user's left ear and / or right ear. In some embodiments, the first detector 120 may include one or more air conduction microphones (which may also be referred to as feedback microphones). For example, the first detector 120 may include a first sub-microphone (or microphone array) and a second sub-microphone (or microphone array). The first sub-microphone (or microphone array) may be located at the user's left ear, and the second sub-microphone (or microphone array) may be located at the user's right ear. The first sub-microphone (or microphone array) and the second sub-microphone (or microphone array) may be simultaneously put into an operating state, or may be controlled so that only one of them is put into an operating state.
[0047] In some embodiments, depending on the operating principle of the microphone, the first detector 120 may include a moving coil microphone, a ribbon microphone, a condenser microphone, an electret microphone, an electromagnetic microphone, a carbon microphone, etc., or any combination thereof. In some embodiments, the arrangement type of the first detector 120 may include a linear array (e.g., straight line, curved), a planar array (e.g., regular and / or irregular shapes such as a cross, circle, ring, polygon, or net), a three-dimensional array (e.g., cylindrical, spherical, hemispherical, polyhedral, etc.), etc., or any combination thereof.
[0048] The processor 130 may be configured to estimate a noise reduction signal for the sound production unit 110 based on an external noise signal, so that the noise reduction signal emitted by the sound production unit 110 reduces or cancels out environmental noise heard by the user, thereby achieving active noise reduction. Specifically, the processor 130 may estimate a second residual signal at a target spatial position based on a first audio signal generated by the sound production unit 110 and a first residual signal picked up by the first detector 120 (including a residual noise signal obtained by superimposing the environmental noise and the first audio signal in the first detector 120). The processor 130 may further update a noise reduction control signal for controlling the sound production by the sound production unit 110 based on the second residual signal. The sound production unit 110 generates a new noise reduction signal in response to the updated noise reduction control signal, thereby achieving real-time modification of the noise reduction signal and a high active noise reduction effect.
[0049] In the present disclosure, the target spatial position may refer to a spatial position within a specific distance from the user's eardrum. The target spatial position may be closer to the user's ear canal (e.g., eardrum) than the first detector 120. The specific distance here may be a fixed distance, such as 0 cm, 0.5 cm, 1 cm, 2 cm, or 3 cm. In some embodiments, the target spatial position may be inside or outside the ear canal. For example, the target spatial position may be the ear membrane, the basilar membrane, or another position outside the ear canal. In some embodiments, the number of microphones in the first detector 120 and their distribution relative to the user's ear canal may be related to the target spatial position. Based on the target spatial position, the number of microphones in the first detector 120 and / or their distribution relative to the user's ear canal may be adjusted. For example, if the target spatial position is closer to the user's ear canal, the number of microphones in the first detector 120 may be increased. As another example, when the target spatial position is closer to the user's ear canal, the spacing between the microphones in the first detector 120 may be reduced. As yet another example, when the target spatial position is closer to the user's ear canal, the arrangement of the microphones in the first detector 120 may be changed.
[0050] In some embodiments, the processor 130 may obtain a first transfer function between the sound production unit 110 and the first detector 120, a second transfer function between the sound production unit 110 and the target spatial location, a third transfer function between the environmental noise source and the first detector 120, and a fourth transfer function between the environmental noise source and the target spatial location. The processor 130 may estimate the second residual signal at the target spatial location based on the first transfer function, the second transfer function, the third transfer function, the fourth transfer function, the first audio signal, and the first residual signal. In some embodiments, the processor 130 may determine the second residual signal by simply obtaining the ratio between the fourth transfer function and the third transfer function, without needing to obtain the third transfer function and the fourth transfer function. In such a case, the processor 130 may obtain a first transfer function between the sound production unit 110 and the first detector 120, a second transfer function between the sound production unit 110 and the target spatial location, and a fifth transfer function (e.g., a ratio of the fourth transfer function to the third transfer function) reflecting the relationship between the environmental noise source, the first detector 120, and the target spatial location. The processor 130 may estimate the second residual signal at the target spatial location based on the first transfer function, the second transfer function, the fifth transfer function, the first audio signal, and the first residual signal. In some embodiments, the processor 130 may obtain only the first transfer function between the sound production unit 110 and the first detector 120 and estimate the second residual signal at the target spatial location based on the first transfer function, the first audio signal, and the first residual signal. For more details on how processor 130 estimates the second residual signal at the target spatial position, reference can be made to other locations in this description (e.g., the portion of FIG. 3 and related discussion), and detailed description will not be given here.
[0051] In some embodiments, processor 130 may include hardware and software modules. By way of example only, the hardware modules may include Digital Signal Processor (DSP) chips and Advanced Reduced Instruction Set Computer Machines (ARM), and the software modules may include algorithmic modules.
[0052] In some embodiments, the acoustic device 100 may further include one or more third detectors (not shown). In some embodiments, the third detector may be referred to as a feedforward microphone. The third detector may be farther from the target spatial location than the first detector 120; i.e., the feedforward microphone is closer to the noise source than the feedback microphone. The third detector may be configured to pick up environmental noise transmitted to the third detector, convert the picked-up environmental noise into an electrical signal, and transmit the electrical signal to the processor 130 for processing. The processor 130 may determine a noise reduction control signal based on the environmental noise acquired by the third detector and the estimated signal at the target spatial location. Specifically, the processor 130 may receive the electrical signal converted from the environmental noise transmitted by the third detector and process it to estimate the environmental noise signal (e.g., the amplitude and phase of the noise) at the target spatial location. The processor 130 may further generate a noise reduction control signal based on the estimated noise signal at the target spatial location. Further, the processor 130 may send a noise reduction control signal to the sound production unit 110. The sound production unit 110 may generate a new noise reduction signal in response to the noise reduction control signal. Parameters (e.g., amplitude, phase, etc.) of the noise reduction signal may correspond to parameters of the environmental noise. By way of example only, the noise reduction signal may have an amplitude approximately equal to the amplitude of the environmental noise and a phase approximately opposite to the phase of the environmental noise, thereby ensuring that the noise reduction signal emitted by the sound production unit 110 has a high active noise reduction effect.
[0053] In some embodiments, the third detector may be located at the left ear and / or right ear of the user. For example, there may be one third detector, and when the user is using the acoustic device 100, the third detector may be located at the user's left ear. As another example, there may be multiple third detectors, and when the user is using the acoustic device 100, the third detectors may be distributed at the user's left ear and right ear so that the acoustic device 100 can better pick up spatial noise transmitted from different sides. In some embodiments, the third detectors may be distributed at various positions on the acoustic device 100, and when the user is using the acoustic device 100, multiple third detectors may be located at the user's left ear, right ear, or may be arranged to surround the user's head.
[0054] In some embodiments, the third detector may be placed in a target area such that an interference signal received from the sound production unit 110 is minimized. If the sound production unit 110 is a bone conduction speaker, the interference signal may include a sound leakage signal and a vibration signal of the bone conduction speaker, and the target area may be an area where the total energy of the sound leakage signal and the vibration signal of the bone conduction speaker transmitted to the third detector is minimized. If the sound production unit 110 is an air conduction speaker, the target area may be an area where the sound pressure level of the radiated sound field of the air conduction speaker is minimized.
[0055] In some embodiments, the third detector may include one or more air conduction microphones. For example, when a user listens to music using the acoustic device 100, the air conduction microphone may simultaneously capture external environmental noise and the user's speech, and the captured external environmental noise and the user's speech may be considered as environmental noise. In some embodiments, the third detector may include one or more bone conduction microphones. The bone conduction microphone is in direct contact with the user's skin, and vibration signals generated by the skeleton or muscles when the user speaks are directly transmitted to the bone conduction microphone. The bone conduction microphone may then convert the vibration signals into electrical signals and transmit the electrical signals to the processor 130 for processing. In some embodiments, the bone conduction microphone is not in direct contact with the human body, and vibration signals generated by the skeleton or muscles when the user speaks are transmitted to a housing structure of the acoustic device 100, and then transmitted to the bone conduction microphone by the housing structure. In some embodiments, when a user is in a call state, the processor 130 can use the environmental noise as the environmental noise to perform noise reduction on the audio signal collected by the air conduction microphone, and transmit the audio signal collected by the bone conduction microphone to the terminal device as a voice signal, thereby ensuring the call quality during the user call (i.e., the quality of the speech of the current user of the audio device 100 and the other party talking to the current user).
[0056] In some embodiments, the processor 130 may control the switch states of the bone conduction microphone and / or the air conduction microphone in the third detector based on the operating state of the acoustic device 100. The operating state of the acoustic device 100 may refer to the usage state when the user is wearing the acoustic device 100. By way of example only, the operating state of the acoustic device 100 may include, but is not limited to, a call state, a non-call state (e.g., a music playback state), a voice message transmission state, etc. In some embodiments, when the third detector picks up environmental noise and a voice signal, the switch states of the bone conduction microphone and the air conduction microphone in the third detector may be determined according to the operating state of the acoustic device 100. For example, when a user wears the acoustic device 100 and plays music, the switch state of the bone conduction microphone may be in a standby state, and the switch state of the air conduction microphone may be in an active state. As another example, when a user wears the acoustic device 100 and sends a voice message, the switch state of the bone conduction microphone may be in an active state, and the switch state of the air conduction microphone may be in an active state. In some embodiments, the processor 130 may control the switch state of a microphone (eg, a bone conduction microphone, an air conduction microphone) in the third detector by sending a control signal.
[0057] In some embodiments, when the operating state of the acoustic device 100 is a non-call state (e.g., a music playback state), the processor 130 may control the bone conduction microphone in the third detector to a standby state and the air conduction microphone to an active state. When the acoustic device 100 is in a non-call state, the user's voice signal may be considered as environmental noise. In this case, the user's voice signal included in the environmental noise picked up by the air conduction microphone may not need to be filtered so that it is canceled out as part of the environmental noise and the noise reduction signal output by the sound output unit 110. When the operating state of the acoustic device 100 is a call state, the processor 130 may control both the bone conduction microphone and the air conduction microphone in the third detector to an active state. When the acoustic device 100 is in a call state, the user's voice signal needs to be put on hold. In this case, the processor 130 sends a control signal to control the bone conduction microphone to an active state, so that the bone conduction microphone can pick up the user's voice signal. The processor 130 removes the user's speech signal picked up by the bone conduction microphone from the environmental noise picked up by the air conduction microphone so that the user's own speech signal is not cancelled out by the noise reduction signal output by the pronunciation unit 110, thereby ensuring the user's normal communication state.
[0058] In some embodiments, when the operating state of the acoustic device 100 is a call state, if the sound pressure of the environmental noise is greater than a preset threshold, the processor 130 may control the bone conduction microphone in the third detector to remain in an operating state. The sound pressure of the environmental noise may reflect the intensity of the environmental noise. The preset threshold may be any value, such as 50 dB, 60 dB, or 70 dB, pre-stored in the acoustic device 100. When the sound pressure of the environmental noise is greater than the preset threshold, the environmental noise may affect the user's call quality. The processor 130 may control the bone conduction microphone to remain in an operating state by sending a control signal. The bone conduction microphone can capture vibration signals of facial muscles when the user speaks while picking up almost no external environmental noise. In this case, the vibration signals picked up by the bone conduction microphone are used as voice signals during a call, ensuring the user's normal call.
[0059] In some embodiments, when the operating state of the acoustic device 100 is a call state, if the sound pressure of the environmental noise is lower than a preset threshold, the processor 130 may control the bone conduction microphone to switch from the operating state to a standby state. When the sound pressure of the environmental noise is lower than the preset threshold, the sound pressure of the environmental noise is lower than the sound pressure of the audio signal generated by the user's speech. In this case, after the user's speech transmitted to the user's ear via the first acoustic path is partially canceled out by the noise reduction signal output by the sound generation unit 110 and transmitted to the user's ear via the second acoustic path, the remaining user's speech is still sufficient to ensure the user's normal call (for example, the user's speech canceled out by the noise reduction signal can be used as a voice signal for the call, converted into an electrical signal, and transmitted to the other acoustic device, where it can be converted into an audio signal by the sound generation unit of the acoustic device, allowing the other user during the call to hear the local user's speech). In this case, the processor 130 controls the bone conduction microphone in the third detector to switch from an active state to a standby state by sending a control signal, thereby further reducing the complexity of signal processing and the power loss of the acoustic device 100. If the sound production unit 110 is an air conduction speaker, the specific position where the noise reduction signal and the environmental noise cancel each other out may be the user's ear canal or its vicinity, such as the eardrum position (i.e., the target spatial position). The first acoustic path may be the path along which the environmental noise is transmitted from the noise source to the target spatial position, and the second acoustic path may be the path along which the noise reduction signal is transmitted from the air conduction speaker to the target spatial position through the air. If the sound production unit 110 is a bone conduction speaker, the specific position where the noise reduction signal and the environmental noise cancel each other out may be the user's basilar membrane. The first acoustic path may be the path along which environmental noise is transmitted from the noise source through the user's ear canal and eardrum to the user's basilar membrane, and the second acoustic path may be the path along which the noise-reducing signal is transmitted from the bone conduction speaker through the user's skeleton or tissue to the user's basilar membrane.
[0060] In some embodiments, the acoustic device 100 may include one or more sensors 140. The one or more sensors 140 may be electrically connected to other components of the acoustic device 100 (e.g., the processor 130). The one or more sensors 140 may acquire physical position and / or motion information of the acoustic device 100. By way of example only, the one or more sensors 140 may include an inertial measurement unit (IMU), a global position system (GPS), radar, etc. The motion information may include a motion trajectory, a motion direction, a motion speed, a motion acceleration, a motion angular velocity, time information related to the motion (e.g., a motion start time, an end time), etc., or any combination thereof. For example, an IMU may include a microelectromechanical system (MEMS). The microelectromechanical system may include a multi-axis accelerometer, a gyroscope, a magnetometer, etc., or any combination thereof. The IMU may detect the physical position and / or movement information of the acoustic device 100 in order to control the acoustic device 100 based on the physical position and / or movement information.
[0061] In some embodiments, the one or more sensors 140 may include a distance sensor. The distance sensor detects the distance from the acoustic device 100 to the user's ear (e.g., the distance between the sound generation unit 110 and the target spatial position), determines the current wearing posture or usage scene of the acoustic device 100 based on the distance, and may further determine a transfer function between the sound generation unit 110, the first detector 120, and the target spatial position. For more details on determining the transfer function based on the distance, please refer to FIG. 3 or FIG. 4 and the description thereof, and the description thereof will be omitted here.
[0062] In some embodiments, the acoustic device 100 may include a memory 150. The memory 150 may store data, instructions, and / or any other information. For example, the memory 150 may store transfer functions between the sound production unit 110, the first detector 120, and target spatial positions corresponding to different users and / or different wearing postures. As another example, the memory 150 may store mapping relationships between transfer functions between the sound production unit 110, the first detector 120, and target spatial positions corresponding to different users and / or different wearing postures. As yet another example, the memory 150 may store data and / or computer programs used to implement the flow 300 shown in FIG. 3 . As yet another example, the memory 150 may store a trained neural network. Note that different users have different tissue morphologies (e.g., different head sizes, different configurations of human body tissues such as muscle tissue, fat tissue, and bone structure), and the corresponding first transfer function, second transfer function, third transfer function, and fourth transfer function may be different. Different wearing postures mean that the wearing position when the user wears the acoustic device 100, the wearing direction of the acoustic device 100, the acting force between the acoustic device 100 and the user, etc. are different, and the corresponding first transfer function, second transfer function, third transfer function and fourth transfer function may also be different.
[0063] In some embodiments, the memory 150 may include mass memory, removable memory, volatile read-write memory, read-only memory (ROM), etc., or any combination thereof. The memory 150 may be signal-connected to the processor 130. When a user wears the acoustic device 100, the processor 130 may obtain corresponding first, second, third, and fourth transfer functions from the memory 150 based on the user's tissue morphology, wearing posture, etc. The processor 130 may estimate a second residual signal at a target spatial position (e.g., eardrum) based on the corresponding first, second, third, and fourth transfer functions to generate a more accurate noise reduction control signal, so that the inverse sound waves emitted by the sound generation unit 110 in response to the noise reduction control signal have a higher active noise reduction effect.
[0064] In some embodiments, the acoustic device 100 may include a signal transceiver 160. The signal transceiver 160 may be electrically connected to other components of the acoustic device 100 (e.g., the processor 130). In some embodiments, the signal transceiver 160 may include Bluetooth®, an antenna, or the like. The acoustic device 100 may communicate with other external devices (e.g., a mobile phone, a tablet personal computer, a smartwatch) via the signal transceiver 160. For example, the acoustic device 100 may wirelessly communicate with other devices via Bluetooth®.
[0065] In some embodiments, the acoustic device 100 may include a housing structure 170. The housing structure 170 may be configured to house other components of the acoustic device 100 (e.g., the sound generation unit 110, the first detector 120, the processor 130, the distance sensor 140, the memory 150, the signal transceiver 160, etc.). In some embodiments, the housing structure 170 may be a hollow, sealed or semi-sealed structure in which other components of the acoustic device 100 are located within or on the housing structure. In some embodiments, the shape of the housing structure 170 may be a three-dimensional structure having a regular or irregular shape, such as a rectangular parallelepiped, a cylinder, or a truncated cone. When the acoustic device 100 is worn by a user, the housing structure 170 may be located near the user's ear. For example, the housing structure 170 may be located around the user's ear (e.g., in front of or behind the user's ear). As another example, the housing structure 170 may be located at the user's ear so as not to block or cover the user's ear canal. In some embodiments, acoustic device 100 is a bone conduction earphone, and at least one side of the housing structure may be in contact with the user's skin. An acoustic driver (e.g., a vibration speaker) in the bone conduction earphone converts audio signals into mechanical vibrations, which can be transmitted to the user's auditory nerves via the housing structure and the user's bones. In some embodiments, acoustic device 100 is an air conduction earphone, and at least one side of the housing structure may or may not be in contact with the user's skin. At least one sound guide hole is included in the side wall of the housing structure, and a speaker in the air conduction earphone converts audio signals into air-conducted sound, which may be emitted toward the user's ear through the sound guide hole.
[0066] In some embodiments, acoustic device 100 may include a securing structure 180. The securing structure 180 may be configured to secure acoustic device 100 in a position near a user's ear without blocking the user's ear canal. In some embodiments, the securing structure 180 may be physically connected (e.g., by fastening or screw connection) to a housing structure 170 of acoustic device 100. In some embodiments, the housing structure 170 of acoustic device 100 may be part of the securing structure 180. In some embodiments, the securing structure 180 may include an ear hook, a back hook, an elastic band, temples, etc. to better secure acoustic device 100 near a user's ear and prevent it from falling off during use by the user. For example, the securing structure 180 may be an ear hook, which may be configured to be worn around the ear region. In some embodiments, the ear hook may be a continuous hook that is elastically pulled to attach to the user's ear, and at the same time, it may apply pressure to the user's pinna, thereby firmly securing acoustic device 100 in a specific position on the user's ear or head. In some embodiments, the ear hook may be a discontinuous strip. For example, the ear hook may include a rigid portion and a flexible portion. The rigid portion may be made of a rigid material (e.g., plastic or metal) and secured to the housing structure 170 of the acoustic device 100 by a physical connection (e.g., a clasp, a screw connection, etc.). The flexible portion may be made of an elastic material (e.g., fabric, a composite material, and / or chloroprene rubber). As another example, the securing structure 180 may be a neckband configured to be worn around the neck / shoulder region. For another example, the securing structure 180 may be temples that are worn over the user's ears as part of eyeglasses.
[0067] In some embodiments, the acoustic device 100 may further include an interactive module (not shown) that adjusts the sound pressure of the noise reduction signal. In some embodiments, the interactive module may include a button, a voice assistant, a gesture sensor, etc. A user can adjust the noise reduction mode of the acoustic device 100 by controlling the interactive module. Specifically, the user can control the interactive module to adjust (e.g., amplify or attenuate) the amplitude information of the noise reduction signal, thereby changing the sound pressure of the noise reduction signal emitted by the sound unit 110 and achieving different noise reduction effects. By way of example only, the noise reduction mode may include a strong noise reduction mode, a medium noise reduction mode, a weak noise reduction mode, etc. For example, when a user is wearing the acoustic device 100 indoors and the external environmental noise is low, the user can use the interactive module to turn off the noise reduction mode of the acoustic device 100 or adjust it to a weak noise reduction mode. As another example, when a user wears acoustic device 100 while walking in a public place such as a street, the user needs to respond to an unexpected situation by listening to audio signals (e.g., music, voice information) while maintaining a certain level of sensitivity to the surrounding environment. In this case, the user can select a moderate noise reduction mode through an interactive module (e.g., a button or a voice assistant) to retain ambient noise (e.g., warning sounds, collision sounds, car horns, etc.). Furthermore, for example, when the user is riding in a vehicle such as a subway or an airplane, the user can select a strong noise reduction mode through the interactive module to further reduce ambient noise. In some embodiments, processor 130 can further alert the user to adjust the noise reduction mode based on the intensity range of the ambient noise by transmitting notification information to acoustic device 100 or a terminal device (e.g., a mobile phone, a smart watch, etc.) communicatively connected to acoustic device 100.
[0068] The above description of FIG. 1 is for illustrative purposes only and is not intended to limit the scope of the present disclosure. Those skilled in the art may make various changes and modifications based on the description of the present disclosure. In some embodiments, one or more components of the acoustic device 100 (e.g., the distance sensor 140, the signal transceiver 160, the fixed structure 180, the interactive module, etc.) may be omitted. In some embodiments, one or more components of the acoustic device 100 may be replaced with other elements capable of achieving similar functions. For example, the acoustic device 100 may not include the fixed structure 180. The housing structure 170 or a portion thereof may have a shape that matches the human ear (e.g., a circular, elliptical, polygonal (regular or irregular), U-shaped, V-shaped, or semicircular) so that the housing structure can be worn near the user's ear. In some embodiments, one component of the acoustic device 100 may be divided into multiple subcomponents, or multiple components may be integrated into a single component. These changes and modifications do not depart from the scope of the present disclosure.
[0069] 2 is a schematic diagram of an acoustic device being worn according to some embodiments of the present disclosure. As shown in FIG. 2, when a user is wearing the acoustic device 200, the acoustic device 200 may be fixed in a position near the user's ear 230 (or head) and not blocking the user's ear canal. The acoustic device 200 may include a sound generation unit 210 and a first detector 220.
[0070] In some embodiments, the first detector 220 may be located on the side of the sound production unit 210 facing the user's ear canal. In some embodiments, the ratio of the acoustic path from the first detector 220 to the target spatial position A to the acoustic path from the first detector 220 to the sound production unit 210 may be between 0.5 and 20. In some embodiments, the acoustic path between the first detector 220 and the target spatial position A may be between 5 mm and 50 mm. In some embodiments, the acoustic path between the first detector 220 and the target spatial position A may be between 15 mm and 40 mm. In some embodiments, the acoustic path between the first detector 220 and the target spatial position A may be between 25 mm and 35 mm. In some embodiments, the number of microphones in the first detector 220 and / or their distribution relative to the user's ear canal may be adjusted based on the acoustic path between the first detector 220 and the target spatial position A.
[0071] Because the acoustic device 200 is an open-type acoustic device (e.g., open-type earphones), the environment in which the first detector 220 and the target spatial position A (e.g., a position close to the user's ear canal and at a specific distance from the eardrum) are located is not a pressure field environment. Therefore, the first detector 220 cannot receive a signal that is completely equal to the signal at the target spatial position A. In this case, by obtaining the correspondence between the audio signal at the first detector 220 and the audio signal at the target spatial position A and then determining the audio signal at the target spatial position A, noise can be reduced more accurately for the target spatial position A.
[0072] 2 is merely an illustrative diagram of the mounting state of the acoustic device, and in embodiments of the present disclosure, the relative positional relationship between the first detector 220, the target spatial position A, and the sound production unit 210 is not limited to that shown in FIG. 2. For example, in some embodiments, the sound production unit 210, the first detector 220, and the target spatial position A do not need to be on the same line. Also, for example, in some embodiments, the first detector 220 may be located away from the target spatial position A of the sound production unit 210, and the distance from the first detector 220 to the target spatial position A may be greater than the distance from the sound production unit 210 to the target spatial position A.
[0073] 3 is a flowchart of an exemplary noise reduction method for an audio device according to some embodiments of the present disclosure. In some embodiments, flow 300 may be performed by audio device 100.
[0074] In step 310, the processor may obtain the first audio signal generated by the sound generation unit 110 based on the noise reduction control signal. In some embodiments, step 310 may be performed by the processor 130.
[0075] In some embodiments, the noise reduction control signal may be generated based on environmental noise picked up by a third detector (i.e., a feedforward microphone). The processor 130 may generate a noise reduction electrical signal (including information in the first audio signal) based on the environmental noise picked up by the third detector, and generate the noise reduction control signal based on the noise reduction electrical signal. Furthermore, the processor 130 may transmit the noise reduction control signal to the sound production unit 110 to cause it to generate the first audio signal. Note that the processor 130 obtaining the first audio signal may also be understood as the processor 130 obtaining the noise reduction electrical signal. The noise reduction electrical signal and the first audio signal differ only in expression; the former is an electrical signal, and the latter is a vibration signal. In some embodiments, the sound production unit 110 may further generate an updated first audio signal based on the updated noise reduction control signal.
[0076] In step 320, the processor may obtain a first residual signal picked up by the first detector 120. The first residual signal may include a residual noise signal obtained by superimposing the environmental noise and the first audio signal in the first detector 120. In some embodiments, step 320 may be performed by the processor 130.
[0077] As can be seen from the related description of FIG. 1 above, environmental noise may refer to a combination of multiple types of external sounds (e.g., traffic noise, industrial noise, construction noise, social noise) in the environment where the user is located. In some embodiments, the first detector 120 may be placed near the user's ear canal to pick up the first residual signal transmitted to the user's ear canal. Furthermore, the first detector 120 may convert the picked-up first residual signal into an electrical signal and transmit it to the processor 130 for processing.
[0078] In step 330, the processor may estimate a second residual signal at a target spatial location based on the first audio signal and the first residual signal. In some embodiments, step 330 may be performed by processor 130.
[0079] The second residual signal may include a residual noise signal in which environmental noise and the first audio signal are superimposed at the target spatial position. Note that because the acoustic device 100 is an open-type acoustic device and the environment in which the first detector 120 (i.e., the feedback microphone) and the target spatial position (e.g., the eardrum) are located is not a pressure field environment, the noise signal received by the first detector 120 cannot directly reflect the noise signal at the target spatial position. Therefore, the processor 130 may determine the second residual signal based on at least one transfer function between the sound generation unit 110, the first detector 120, the environmental noise source, and the target spatial position. In some embodiments, the transfer function between any two of the sound generation unit 110, the first detector 120, the environmental noise source, and the target spatial position may represent a relationship between the audio signals at the corresponding positions of the two, for example, reflecting the transmission quality in the process of transmitting an audio signal generated by one to the other, or the relationship between an audio signal acquired by one and an audio signal generated by the other. For example, the transfer function between the sound producing unit 110 and the first detector 120 may represent the transmission quality in a transmission process in which the first audio signal generated by the sound producing unit 110 is transmitted to the first detector 120, or the relationship between the first residual signal picked up by the first detector 120 and the first audio signal generated by the sound producing unit 110. As another example, the transfer function between the environmental noise source and the first detector 120 may represent the transmission quality in a transmission process in which environmental noise is transmitted from the environmental noise source to the first detector 120, or the relationship between the first residual signal picked up by the first detector 120 and the environmental noise generated by the environmental noise source.
[0080] In some embodiments, the first audio signal (also called the noise-reduced signal) emitted by the sound output unit 110 may be S, and the environmental noise signal may be N. In this case, the signal at the first detector 120 (i.e., the first residual signal) M and the signal at the target spatial position (i.e., the second residual signal) D can be expressed by equations (1) and (2), respectively.
[0081]
number
[0082] where H SM represents the first transfer function between the sound generation unit 110 and the first detector 120, and H SD represents the second transfer function between the sound generation unit 110 and the target spatial position, and H NM represents a third transfer function between the environmental noise source and the first detector 120, and H ND represents the fourth transfer function between the environmental noise source and the target spatial position.
[0083] To achieve the goal of active noise reduction, it is necessary to estimate the second residual signal D at the target spatial position. The second residual signal D at the target spatial position can be regarded as the noise volume audible to the user after active noise reduction (e.g., the signal that can be received by the user's eardrum). In this case, the above equations (1) and (2) can be simplified to the following equation (3):
[0084]
number
[0085] In some embodiments, the processor 130 calculates a first transfer function H between the sound generation unit 110 and the first detector 120. SM , the second transfer function H between the sound generation unit 110 and the target spatial position SD , a third transfer function H between the environmental noise source and the first detector 120 NM , and a fourth transfer function H between the environmental noise source and the target spatial position. NDmay be directly obtained. Furthermore, the processor 130 may estimate the second residual signal D at the target spatial position based on the first transfer function, the second transfer function, the third transfer function, the fourth transfer function, the first audio signal S, and the first residual signal M according to Equation (3). In some embodiments, the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function may be associated with a user category. The processor 130 may directly retrieve the corresponding first transfer function, the second transfer function, the third transfer function, and the fourth transfer function from the memory 150 based on the current user category (e.g., adult or child).
[0086] In some embodiments, the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function may be related to the wearing position of the acoustic device 100. The processor 130 may directly call up the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function corresponding to the current wearing position from the memory 150. For example, the acoustic device 100 may include one or more sensors, such as a distance sensor or a position sensor. The sensor may detect the distance from the acoustic device 100 to the user's ear and / or the relative position of the acoustic device 100 and the user's ear. Different wearing positions of the acoustic device 100 may correspond to different distances from the acoustic device 100 to the user's ear and / or different relative positions of the acoustic device 100 and the user's ear. The processor 130 can determine the current wearing position of the acoustic device 100 based on the distance data and / or position data acquired by the sensor, and further determine a first transfer function, a second transfer function, a third transfer function and a fourth transfer function corresponding to the current wearing position.
[0087] In some embodiments, the processor 130 may directly determine the first, second, third, and fourth transfer functions corresponding to the acoustic device 100 based on sensing data from a sensor (e.g., the relative positional relationship or distance relationship between the acoustic device 100 and the user's ear, etc.). Specifically, different distances from the acoustic device 100 to the user's ear and / or different relative positions between the acoustic device 100 and the user's ear may correspond to different first, second, third, and fourth transfer functions. The processor 130 may directly call the first, second, third, and fourth transfer functions corresponding to the distance data and / or position data acquired by the sensor.
[0088] In some embodiments, the first transfer function may have a mapping relationship with the second transfer function, the third transfer function, and the fourth transfer function, respectively. The processor 130 may acquire the first transfer function and determine the second transfer function, the third transfer function, and the fourth transfer function, respectively, based on the mapping relationship between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function, thereby determining the second residual signal D at the target spatial position. In some embodiments, the mapping relationship between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function may be determined by a trained neural network. Specifically, the processor 130 may determine the first transfer function between the sound generation unit 110 and the first detector 120 based on the relationship between the first audio signal (the noise control signal for generating the first audio signal) and the first residual signal. For example, when a user is wearing the acoustic device 100, in a noise-free situation, the first transfer function may be determined by the following equation (4):
[0089]
number
[0090] Furthermore, the processor 130 may input the first transfer function into a trained neural network and obtain the output of the trained neural network to obtain the second transfer function, the third transfer function, and / or the fourth transfer function.
[0091] In some embodiments, the mapping relationships between the first transfer function and each of the second, third, and fourth transfer functions may be generated based on test data for different wearing scenes (or different wearing postures) of the acoustic device 100 and stored in the memory 150. The processor 130 can directly call and use them. Note that in different wearing scenes or usage states, the acoustic device 100 may correspond to different first, second, third, and fourth transfer functions. Furthermore, the first transfer function and the second, third, and fourth transfer functions may have different mapping relationships, and these mapping relationships may change with changes in the wearing scene (or wearing posture), etc. For more details about the mapping relationships between the first transfer function and each of the second, third, and fourth transfer functions, please refer to FIG. 4 and its description; a detailed description thereof will be omitted here.
[0092] In some embodiments, the processor 130 may determine a relationship between the second residual signal and the first transfer function, the first audio signal, and the first residual signal based on a mapping relationship between the first transfer function and each of the second, third, and fourth transfer functions. In other words, the second residual signal can be regarded as a function with the first transfer function as a variable. After determining the first transfer function, the processor 130 can estimate the second residual signal at the target spatial position based on the first transfer function, the first audio signal generated by the sound generation unit 110, and the first residual signal received by the first detector 120.
[0093] TIFF0007719546000004.tif65161
[0094] In some embodiments, the second transfer function and the first transfer function may have a first mapping relationship, and the fifth transfer function and the first transfer function may have a second mapping relationship. After determining the first transfer function, the processor 130 may determine the second transfer function based on the first transfer function and the first mapping relationship between the first transfer function and the second transfer function, and may determine the fifth transfer function (i.e., the ratio between the fourth transfer function and the third transfer function) based on the ratio between the fourth transfer function and the third transfer function and the second mapping relationship between the first transfer function. For more information about the first mapping relationship and the second mapping relationship, please refer to FIG. 4 and its description, and the description will be omitted here.
[0095] In some embodiments, the acoustic device 100 may further include an adjustment button or may be adjusted by an application program (APP) on a user terminal. The adjustment button or the APP on the user terminal allows a user to select a transfer function or a mapping relationship between transfer functions associated with the acoustic device 100 as needed by the user. For example, the user may select a distance from the acoustic device 100 to the user's ear (or face) (i.e., adjust the wearing posture) by using the adjustment button or the APP on the user terminal. The processor 130 may obtain a corresponding first transfer function, a second transfer function, a third transfer function, and a fourth transfer function, or a mapping relationship between the first transfer function, the second transfer function, the third transfer function, and / or the fourth transfer function, based on the distance from the acoustic device 100 to the user's ear (or face). Furthermore, the processor 130 may estimate a second residual signal D at a target spatial position based on the obtained transfer function or the mapping relationship between the transfer functions, the first audio signal S of the sound generation unit 110, and the first residual signal M detected by the first detector 120. In other words, the user can adjust the active noise reduction performance of the acoustic device 100, for example, full noise reduction or partial noise reduction, through the adjustment button or the APP on the user terminal.
[0096] In step 340, the processor may update the noise control signal of the sound generation unit 110 based on the second residual signal at the target spatial location. In some embodiments, step 340 may be performed by the processor 130.
[0097] In some embodiments, the processor 130 may generate a corresponding new noise-reduction electrical signal based on the second residual signal D estimated in step 330, and may generate a new noise-reduction control signal based on the new noise-reduction electrical signal. Alternatively, the processor 130 may update the noise-reduction control signal that controls the sound generation unit 110 to generate sound. Specifically, in some embodiments, when complete active noise reduction needs to be achieved, the second residual signal D at the target spatial position can be regarded as essentially 0, that is, the acoustic device 100 can essentially remove external noise, preventing the user from hearing the external noise, and achieving a high active noise reduction effect. In this case, the first sound signal S emitted by the sound generation unit 110 may be simplified as follows:
[0098]
number
[0099] In other words, the processor 130 calculates the first transfer function H between the sound generation unit 110 and the first detector 120. SM , the second transfer function H between the sound generation unit 110 and the target spatial position SD , a third transfer function H between the environmental noise source and the first detector 120 NM , the fourth transfer function H between the environmental noise source and the target spatial position ND, and the first residual signal M in the first detector 120, the magnitude of the noise reduction signal that needs to be emitted by the sound producing unit 110 is calculated, thereby correcting the noise reduction signal emitted by the conventional sound producing unit 110, realizing real-time correction of the noise reduction signal of the sound producing unit 110, and ensuring that the noise reduction signal emitted by the sound producing unit 110 can achieve a high active noise reduction effect.
[0100] It should be noted that the description of the flow 300 is merely for illustrative and explanatory purposes and does not limit the scope of the present disclosure. Those skilled in the art can make various modifications and changes to the flow 300 based on the description of the present disclosure. These modifications and changes still fall within the scope of the present disclosure. For example, in some embodiments, the acoustic device 100 may be a closed-type acoustic device, i.e., the first detector 120 and the target spatial location may be located in a pressure sound field. In this case, H NM= H ND and H SD= H SM As can be seen from equation (3), the signal M (i.e., the first residual signal) at the first detector 120 and the signal D (i.e., the second residual signal) at the target spatial position are the same. The noise-reduced signal S (i.e., the first audio signal) emitted by the sound generation unit 110 can satisfy the following relationship:
[0101]
number
[0102] In this case, the processor 130 calculates the first transfer function H SM , a third transfer function H between the environmental noise source and the first detector 120 NMBy estimating the noise reduction signal that needs to be emitted by the sound production unit 110 based on the acquired signal M and the environmental noise signal N in the first detector 120, the noise reduction signal emitted by the conventional sound production unit 110 can be corrected, real-time correction of the noise reduction signal can be realized, and a high active noise reduction effect can be achieved.
[0103] In some embodiments, when the acoustic device 100 is a closed-type acoustic device and needs to achieve complete active noise reduction, the second residual signal D at the target spatial position and the first residual signal M at the first detector 120 may be regarded as essentially 0. In this case, the noise-reduced signal S (i.e., the first audio signal) emitted by the sound generation unit 110 can satisfy the following relationship:
[0104]
number
[0105] In this case, the external noise can be completely removed by the noise-reducing signal emitted by the sound-producing unit 110. The processor 130 calculates the first transfer function H SM , a third transfer function H between the environmental noise source and the first detector 120 NM , and the environmental noise signal N, the magnitude of the noise reduction signal that needs to be emitted by the sound production unit 110 is estimated, thereby modifying the noise reduction signal emitted by the conventional sound production unit 110, thereby realizing real-time modification of the noise reduction signal emitted by the sound production unit 110, and ensuring that the noise reduction signal emitted by the sound production unit 110 can achieve a high active noise reduction effect.
[0106] It should be noted that the description of the flow 300 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may make various modifications and changes to the flow 300 based on the description of the present disclosure. These modifications and changes still fall within the scope of the present disclosure. In some embodiments, the flow 300 may be stored in a computer-readable storage medium in the form of computer instructions. When the computer instructions are executed, the noise reduction method described above can be realized.
[0107] FIG. 4 is an exemplary flowchart of a method for determining a transfer function of an acoustic device according to some embodiments of the present disclosure. In some embodiments, the acoustic device includes at least a sound generation unit, a first detector, a processor, and a fixing structure. When a user wears the acoustic device, the fixing structure can fix the acoustic device in a position near the user's ear but not blocking the user's ear canal, such that the target spatial position (e.g., the user's eardrum or basilar membrane) is closer to the user's ear canal than the first detector. For more details about the sound generation unit, the first detector, the processor, the target spatial position, etc., please refer to the description related to the acoustic device 100 in FIG. 1, and therefore, a description thereof will be omitted here. In some embodiments, the steps in flow 400 may be called and / or executed by the processor 130 in the acoustic device 100 or another processing device other than the processor 130.
[0108] In step 410, the processor 130 may obtain a first signal emitted by the sound generation unit 110 based on the noise reduction control signal in a scene without environmental noise, and a second signal picked up by the first detector.
[0109] Specifically, the processor 130 can input a noise reduction control signal to the sound output unit 110 after the subject wears the acoustic device 100. The sound output unit 110 can output a first signal S0 in response to receiving the noise reduction control signal. The first signal S0 output by the sound output unit 110 can then be transmitted to the first detector 120 and picked up by it. Note that the signal M0 (e.g., the second signal) picked up by the first detector 120 may differ from the first signal S0 due to energy loss during the transmission of the first signal, reflections between the signal and the subject and / or the acoustic device 100, and noise in the environment. Note that different subjects have different body tissue morphologies (e.g., different head sizes, different tissue configurations such as muscle tissue, fat tissue, and bone structure), and therefore the wearing posture (e.g., different wearing positions, different contact forces with the subject) of the acoustic device may differ. In some embodiments, even for the same subject, the wearing posture (e.g., wearing position) of the acoustic device 100 may be different. When the wearing posture is different, even if the relative positions of the sound generation unit 110 and the first detector 120 do not change during the process in which the signal emitted by the sound generation unit 100 is transmitted to the first detector 120, the transmission conditions during the transmission process of the signal emitted by the sound generation unit 110 change (e.g., the signal reflection conditions differ) due to the different wearing posture of the subject. Therefore, the first transfer function between the sound generation unit 110 and the first detector 120 of the acoustic device 100 may differ depending on the wearing posture.
[0110] In some embodiments, the subject may be a head model in a laboratory or a user. For example, when the acoustic device 100 is attached to a head model, the first detector 120 and the sound generation unit 110 of the acoustic device 100 may be located near the ear canal of the head model. In some embodiments, the control signal may be an electrical signal including any audio signal. Note that in the present disclosure, audio signals (e.g., first and second signals) may include parameter information such as frequency information, amplitude information, and phase information. In some embodiments, the first and / or second signals may be audio signals or electrical signals obtained by converting audio signals.
[0111] In step 420, the processor 130 may determine a first transfer function between the sound generation unit 110 and the first detector 120 based on the first signal and the second signal.
[0112] It can be understood that in a scene without environmental noise, the second signal M0 detected by the first detector 120 is all transmitted from the sound-producing unit 110. The ratio of the second signal M0 picked up by the first detector 120 to the first signal S0 output by the sound-producing unit 110 can directly reflect the transmission quality or transmission efficiency of the first signal generated by the sound-producing unit 110 during transmission from the sound-producing unit 110 to the first detector 120. In some embodiments, the first transfer function H SM has a positive correlation with the ratio of the second signal M0 to the first signal S0. SM and the first signal S0 and the second signal M0 may satisfy the following relationship:
[0113]
number
[0114] In step 430, the processor 130 may acquire a third signal picked up by the second detector. The second detector may be placed at a target spatial position to simulate the eardrum (or basilar membrane) of a human ear and pick up the sound signal. The target spatial position may be closer to the subject's ear canal than the first detector 120. In some embodiments, the target spatial position may be the subject's ear canal, eardrum, or basilar membrane. For example, if the sound production unit 110 is an air conduction speaker, the target spatial position may be at or near the eardrum of the subject. If the sound production unit 110 is a bone conduction speaker, the target spatial position may be at or near the basilar membrane of the subject. In some embodiments, the second detector may be a micro-microphone (e.g., a MEMS microphone) that can enter the user's ear canal and collect sound inside the ear canal.
[0115] Specifically, the first signal S0 output by the sound output unit 110 may be transmitted to a target spatial position and picked up by a second detector at the target spatial position. Similar to the transmission of the first signal to the first detector 120, the signal D0 (e.g., a third signal) picked up by the second detector may differ from the first signal S0 due to energy loss during the transmission of the first signal, reflections between the signal and the subject and / or the acoustic device 100, noise in the environment, etc. Furthermore, the second transfer function between the sound output unit 110 of the acoustic device 100 and the target spatial position (second detector) may differ depending on the wearing posture.
[0116] In step 440, the processor 130 may determine a second transfer function between the sound producing unit 110 and the target spatial position based on the first signal and the third signal.
[0117] It can be understood that in a scene without environmental noise, the third signal D0 detected by the second detector is all transmitted from the sound-producing unit 110. The ratio of the third signal D0 picked up by the second detector to the first signal S0 output by the sound-producing unit 110 can directly reflect the transmission quality or transmission efficiency of the transmission process in which the first signal generated by the sound-producing unit 110 is transmitted from the sound-producing unit 110 to the second detector (i.e., the target spatial position). In some embodiments, the second transfer function H SD may be positively correlated with the ratio of the third signal D0 to the first signal S0. By way of example only, the second transfer function H SD and the first signal S0 and the third signal D0 may satisfy the following relationship:
[0118]
number
[0119] In step 450, the processor 130 may acquire a fourth signal picked up by the first detector 120 and a fifth signal picked up by the second detector in a scene where there is environmental noise and the sound production unit 110 does not emit any signal. The environmental noise may be generated by one or more environmental noise sources. In the testing process, the environmental noise source may be any sound source other than the sound production unit. For example, the environmental noise N0 may be simulated and acquired by another sound production device in the test environment.
[0120] TIFF0007719546000010.tif54161
[0121] In step 460, the processor 130 may determine a third transfer function between the environmental noise source and the first detector 120 based on the environmental noise and the fourth signal.
[0122] TIFF0007719546000011.tif64161
[0123]
number
[0124] In step 470, processor 130 may determine a fourth transfer function between the environmental noise source and the target spatial location based on the environmental noise and the fifth signal.
[0125] TIFF0007719546000013.tif62160
[0126]
number
[0127] In some embodiments, the memory 150 may store a first transfer function, a second transfer function, a third transfer function, and a fourth transfer function measured for a certain category of subjects (e.g., adults, children). When a user wears the acoustic device 100, the processor 130 may directly call the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function measured for a typical subject to roughly estimate a second residual signal at a target spatial position (e.g., at the user's eardrum) and roughly estimate a noise reduction signal of the sound-producing unit to achieve active noise reduction. For example, a set of the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function may correspond to an adult male. A set of the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function may correspond to a child. If the user is a child, the processor 130 may call a set of the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function corresponding to the child.
[0128] In some embodiments, the processor 130 repeats steps 410 to 470 for different wearing scenes (e.g., different wearing positions) or different subjects to determine multiple sets of transfer functions for different wearing postures of the acoustic device 100. The multiple sets of transfer functions corresponding to the different wearing postures can be stored in the memory 150 for recall. Each set of transfer functions may include a corresponding first transfer function, a second transfer function, a third transfer function, and a fourth transfer function. When a user is wearing the acoustic device 100, the processor 130 can recall the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function corresponding to the wearing posture based on the wearing posture of the acoustic device 100. Furthermore, the processor 130 can estimate a second residual signal at a target spatial position based on the recalled transfer function, the first audio signal of the sound generation unit 110, and the first residual signal picked up by the first detector 120, and update a noise reduction control signal for controlling the sound generation of the sound generation unit 110 based on the second residual signal. For more details about determining the second residual signal based on the transfer function, please refer to FIG. 3 and the description thereof, and the description will be omitted here.
[0129] In some embodiments, since the transfer function changes depending on the wearing posture of the acoustic device 100, when a user is wearing the acoustic device 100, the processor 130 can directly determine the first transfer function based on the first audio signal output by the sound generation unit 110 and the first residual signal detected by the first detector 120, but cannot directly obtain the second transfer function, the third transfer function, and the fourth transfer function. In this case, the processor 130 may determine the second transfer function, the third transfer function, and the fourth transfer function based on the first transfer function and the relationship between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function, respectively. Specifically, the processor 130 can determine the relationship between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function based on multiple sets of transfer functions corresponding to different wearing postures, and store the relationship in the memory 150 for recall. In some embodiments, processor 130 may determine the relationship between the first transfer function and each of the second, third, and fourth transfer functions through statistics. In some embodiments, processor 130 may train the neural network using multiple sets of sample transfer functions as training samples. Each set of sample transfer functions may be actually measured using a test signal in a different wearing state of acoustic device 100. Processor 130 may determine the relationship between the first transfer function and each of the second, third, and fourth transfer functions as the trained neural network. For example, for the relationship between the first transfer function and the second transfer function, processor 130 may train the first neural network using the first sample transfer function in each set of sample transfer functions as the input of the first neural network and the second sample transfer function in the set of sample transfer functions as the output of the first neural network. The processor 130 may train the first neural network to relate the first transfer function to the second transfer function.Specifically, in application, the processor 130 may input the first transfer function into a trained first neural network to determine the second transfer function.
[0130] In some embodiments, the third transfer function H NM and the fourth transfer function H ND may be considered as a whole, in which case the third transfer function H NM and the fourth transfer function H ND In this case, the processor 130 may determine the second residual signal without acquiring the first transfer function H based on multiple sets of transfer functions corresponding to different wearing positions. SM and the second transfer function H SD The first mapping relationship with the third transfer function H NM and the fourth transfer function H ND and the first transfer function H SM and a second mapping relationship between the first and second mapping relationships may be determined, and the first and second mapping relationships may be stored in the memory 150 for future reference. Illustratively, the first and second mapping relationships may be expressed as follows:
[0131]
number
[0132] When a user is wearing the acoustic device 100, the processor 130 can determine a second transfer function based on the first transfer function and the first mapping relationship, and can determine a ratio between the fourth transfer function and the third transfer function based on the first transfer function and the second mapping relationship. Furthermore, the processor 130 can estimate a second residual signal at a target spatial position based on the first transfer function, the second transfer function, the ratio between the fourth transfer function and the third transfer function, the first audio signal emitted by the sound production unit 110, and the first residual signal detected by the first detector 120, and can update the noise control signal based on the second residual signal at the target spatial position. The sound production unit 110 generates a new first audio signal (i.e., a noise-reduced signal) in response to the updated noise control signal.
[0133] In some embodiments, the processor 130 may train a neural network using multiple sets of sample transfer functions as training samples, obtain a trained neural network, and define the trained neural network as the second mapping relationship. Specifically, the processor 130 may train the second neural network by using a first sample transfer function in each set of sample transfer functions as an input to the second neural network and a ratio between a fourth sample transfer function and a third sample transfer function in the set of sample transfer functions as an output of the second neural network. The processor 130 may define the trained second neural network as the second mapping relationship. In application, the processor 130 may input the first transfer function to the trained second neural network to determine the ratio between the fourth transfer function and the third transfer function.
[0134] In some embodiments, the acoustic device 100 may include one or more sensors (which may be referred to as a fourth detector), such as a distance sensor or a position sensor. The sensor may detect the distance between the acoustic device 100 and the user's ear (or face) and / or the relative position of the acoustic device 100 with respect to the user's ear. For ease of explanation, this disclosure will describe the sensor using a distance sensor as an example. In some embodiments, different wearing positions may correspond to different distances between the acoustic device 100 and the user's ear (or face). The processor 130 may store a first transfer function, a second transfer function, a third transfer function, and a fourth transfer function corresponding to different distances in the memory 150 and prepare for recall. In some embodiments, the processor 130 may store different wearing positions of the acoustic device 100 and the corresponding distances and transfer functions in the memory 150. When the user is wearing the acoustic device 100, the processor 130 may first determine the wearing orientation of the acoustic device 100 based on the distance between the acoustic device 100 and the user's ear detected by the distance sensor (i.e., the fourth detector). The processor 130 may further determine the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the wearing orientation. Alternatively, the processor 130 may directly determine the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the distance between the acoustic device 100 and the user's ear detected by the distance sensor (i.e., the fourth detector). In some embodiments, the processor 130 may determine a mapping relationship between the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the distance between the acoustic device 100 and the user's ear detected by the distance sensor.
[0135] In some embodiments, processor 130 may use distance data acquired by the distance sensor (or the distance data together with the first transfer function) as input to a trained third neural network to obtain the second transfer function, the third transfer function, and / or the fourth transfer function. Specifically, processor 130 may use sample distances acquired by the distance sensor (or the sample distances together with a first sample transfer function from a corresponding set of sample transfer functions) as input to the third neural network, and use the second sample transfer function, the third sample transfer function, and / or the fourth sample transfer function from the set of sample transfer functions as output from the third neural network to train the third neural network. In application, processor 130 may input the distance data acquired by the distance sensor (or the distance data together with the first transfer function) to the trained third neural network to determine the second transfer function, the third transfer function, and / or the fourth transfer function.
[0136] The above description of flow 400 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may make various modifications and variations to flow 400 based on the teachings of the present disclosure. These modifications and variations are still within the scope of the present disclosure. For example, in some embodiments, during the testing process, the second signal may be acquired first, the third signal may be acquired first, or the second and third signals may be acquired simultaneously. In some embodiments, flow 400 may be stored in a computer-readable storage medium in the form of computer instructions. When the computer instructions are executed, the above-described transfer function testing method can be realized.
[0137] Although the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure above is merely illustrative and does not limit the present disclosure. Although not explicitly described in the present disclosure, those skilled in the art may make various changes, improvements, and modifications to the present disclosure. These changes, improvements, and modifications are intended to be suggested by the present disclosure and therefore fall within the spirit and scope of the exemplary embodiments of the present disclosure.
[0138] Furthermore, certain terms are used in this disclosure to describe embodiments of the present disclosure. For example, "one embodiment," "one embodiment," and / or "some embodiments" refer to particular features, structures, or characteristics associated with at least one embodiment of the present disclosure. Therefore, it is emphasized and understood that two or more references to "one embodiment" or "one embodiment" or "one alternative embodiment" in various parts of this disclosure do not necessarily all refer to the same embodiment. Furthermore, certain features, structures, or characteristics in one or more embodiments of the present disclosure may be combined as appropriate.
[0139] Additionally, as will be appreciated by those skilled in the art, aspects of the present disclosure may be illustrated and described in several patentable classes or contexts, including any new and useful process, machine, manufacture, or combination of matter, or any new and useful improvement thereto. Accordingly, aspects of the present disclosure may be implemented entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. Such hardware or software may be referred to as a "data block," "module," "engine," "unit," "assembly," or "system." Additionally, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable medium(s) containing computer-readable program code.
[0140] The computer storage medium may include a propagated data signal, propagated in baseband or as part of a carrier wave, for carrying computer program code. The propagated signal may take various forms, such as an electromagnetic signal, an optical signal, or a suitable combination. The computer storage medium may be any computer-readable medium other than a computer-readable storage medium, which can be coupled to an instruction execution system, device, or apparatus to achieve communication, propagation, or transmission of a program used therein. The program code on the computer storage medium may be propagated via any suitable medium, including wireless, cable, fiber optic cable, RF, or similar media, or any combination of the above media.
[0141] Additionally, unless expressly stated in the claims, the enumerated order, use of alphanumeric characters, or use of other designations of processing elements or sequences described in this disclosure do not limit the order of the procedures and methods of the present disclosure. While the above disclosure has described through various examples what are presently believed to be various useful embodiments of the invention, it should be understood that such details are merely illustrative, and that the appended claims are not limited to the disclosed embodiments, but rather are intended to cover all modifications and equivalent combinations within the spirit and scope of the disclosed embodiments. For example, the system assembly described above may be implemented by a hardware device, or may be implemented as a software-only solution, for example, by installing the described system on an existing server or mobile device.
[0142] Similarly, in the foregoing description of embodiments of the present disclosure, it should be understood that various features may be grouped together in a single embodiment, drawing, or description for the purpose of simplifying the disclosure and facilitating an understanding of one or more embodiments of the present disclosure. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed subject matter requires more features than are recited in each claim. In fact, an embodiment may include fewer than all features of a single embodiment disclosed above.
[0143] In some embodiments, numbers describing the number of components and attributes are used; it should be understood that the numbers describing such embodiments are, in some instances, modified by the modifiers "about," "approximately," or "generally." Unless otherwise specified, "about," "approximately," or "generally" indicates that the stated number is subject to a ±20% variation. Thus, in some embodiments, all numerical parameters used in the present disclosure and claims are approximations that may vary depending on the specific characteristics required for a particular embodiment. In some embodiments, numerical parameters should be calculated taking into account the number of significant digits provided and ordinary rounding techniques should be applied. While in some embodiments, the numerical ranges and parameters used to determine ranges are approximations, in specific embodiments, such numerical values are set as precisely as possible.
[0144] All patents, patent applications, published patent applications, and other materials, such as papers, books, specifications, publications, documents, etc., referenced in this disclosure are incorporated by reference in their entirety into this disclosure, except for prosecution history documents that are inconsistent with or inconsistent with the content of this disclosure and documents that may have a limiting effect on the broadest scope of the claims of this disclosure (now or later related to this disclosure). In addition, if the explanations, definitions, and / or term usage in the accompanying materials of this disclosure are inconsistent with or inconsistent with the content set forth in this disclosure, the explanations, definitions, and / or term usage in this disclosure shall control.
[0145] Finally, it should be understood that the embodiments described herein are merely illustrative of the principles of embodiments of the present disclosure. Other variations may be within the scope of the present disclosure. Thus, by way of example, and not of limitation, alternative configurations of the embodiments of the present disclosure may be considered consistent with the teachings of the present disclosure. Thus, the embodiments of the present disclosure are not limited to the embodiments expressly introduced and described herein. [Explanation of symbols]
[0146] 100 Sound equipment 110 pronunciation units 120 First Detector 130 processors 140 sensors 150 memory 160 Signal Transmitter / Receiver 170 Housing Structure 180 Fixed structure 200 Sound equipment 210 Sound Unit 220 First Detector
Claims
1. An acoustic device, comprising: a sound generating unit; a first detector; a processor; and a fixed structure; The sound generating unit generates a first sound signal based on a noise reduction control signal; The first detector picks up a first residual signal including a residual noise signal in which environmental noise and the first audio signal are superimposed at the first detector; the processor estimates a second residual signal at a target spatial location based on the first audio signal and the first residual signal, and updates the noise reduction control signal based on the second residual signal; the fixing structure fixes the acoustic device at a position near the ear of the user and not blocking the ear canal of the user so that the target spatial position is closer to the ear canal of the user than the first detector; estimating a second residual signal at a target spatial location based on the first audio signal and the first residual signal includes: obtaining a first transfer function between the sound producing unit and the first detector, a second transfer function between the sound producing unit and the target spatial position, a third transfer function between an environmental noise source and the first detector, and a fourth transfer function between the environmental noise source and the target spatial position; Obtaining the first transfer function; determining the second transfer function, the third transfer function, and the fourth transfer function based on the first transfer function and mapping relationships between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function, respectively; or inputting the first transfer function into a trained neural network and obtaining outputs of the trained neural network as the second transfer function, the third transfer function, and the fourth transfer function. and estimating the second residual signal at the target spatial position based on the first transfer function, the second transfer function, the third transfer function, the fourth transfer function, the first audio signal, and the first residual signal.
2. The acoustic device according to claim 1 , wherein each mapping relationship between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function is generated based on test data in different wearing scenes of the acoustic device.
3. Obtaining the first transfer function includes: The acoustic device according to claim 1 , further comprising: calculating the first transfer function based on the noise reduction control signal and the first residual signal.
4. a distance sensor for detecting a distance from the acoustic device to the ear of the user; The acoustic device of claim 1 , wherein the processor is further configured to determine the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the distance.
5. The acoustic device according to claim 1 , wherein the target spatial position is a position of the eardrum of the user.
6. An acoustic device, comprising: a sound generating unit; a first detector; a processor; and a fixed structure; The sound generating unit generates a first sound signal based on a noise reduction control signal; The first detector picks up a first residual signal including a residual noise signal in which environmental noise and the first audio signal are superimposed at the first detector; the processor estimates a second residual signal at a target spatial location based on the first audio signal and the first residual signal, and updates the noise reduction control signal based on the second residual signal; the fixing structure fixes the acoustic device at a position near the ear of the user and not blocking the ear canal of the user so that the target spatial position is closer to the ear canal of the user than the first detector; estimating a second residual signal at a target spatial location based on the first audio signal and the first residual signal includes: obtaining a first transfer function between the sound producing unit and the first detector, a second transfer function between the sound producing unit and the target spatial position, and a fifth transfer function reflecting a relationship between an environmental noise source, the first detector, and the target spatial position, wherein the first transfer function and the second transfer function have a first mapping relationship, and the fifth transfer function and the first transfer function have a second mapping relationship; and estimating the second residual signal at the target spatial position based on the first transfer function, the second transfer function, the fifth transfer function, the first audio signal, and the first residual signal.
7. 1. A method for determining a transfer function of an acoustic device, the acoustic device including a sound generating unit, a first detector, a processor, and a fixed structure, the fixed structure fixing the acoustic device in a position near an ear of a subject and not blocking the ear canal of the subject, the method comprising: In a scene without environmental noise, obtaining a first signal emitted by the sound generating unit based on a noise reduction control signal and a second signal picked up by the first detector, the second signal including a residual noise signal transmitted to the first detector by the first signal; determining a first transfer function between the sound generating unit and the first detector based on the first signal and the second signal; acquiring a third signal picked up by a second detector positioned at a target spatial location closer to the subject's ear canal than the first detector, the third signal including a residual noise signal transmitted to the target spatial location by the first signal; determining a second transfer function between the sound-producing unit and the target spatial position based on the first signal and the third signal; obtaining a fourth signal picked up by the first detector and a fifth signal picked up by the second detector in a scene where the environmental noise is present and the sounding unit does not transmit any signal; determining a third transfer function between the environmental noise source and the first detector based on the environmental noise and the fourth signal; determining a fourth transfer function between the environmental noise source and the target spatial location based on the environmental noise and the fifth signal; determining multiple sets of transfer functions for different wearing scenes or different subjects, each set of transfer functions including a corresponding first transfer function, a second transfer function, a third transfer function, and a fourth transfer function; determining a mapping relationship between the first transfer function and the second transfer function, the third transfer function, and the fourth transfer function based on the plurality of sets of transfer functions.
8. determining a mapping relationship among the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the plurality of sets of transfer functions, training a neural network using the sets of transfer functions as training samples; and training a neural network to map relationships between the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function.
9. determining a mapping relationship among the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the plurality of sets of transfer functions, acquiring a distance from the acoustic device to a corresponding ear of the subject for each of the different wearing scenes or the different subjects; and determining a mapping relationship between the first transfer function, the second transfer function, the third transfer function, and the fourth transfer function based on the distance and the sets of transfer functions.
Citation Information
Patent Citations
Active noise control with compensation for error sensing at the eardrum
US20140044275A1
Active noise cancelling systems and methods
US20210304725A1
Acoustic processing device and acoustic processing method
WO2019053993A1