Acoustic device

By generating a noise-reducing signal through sound field estimation using a microphone array and processor, and outputting the target signal using a speaker, the problem of obstruction in in-ear acoustic devices and insufficient noise reduction in open-back acoustic devices is solved, achieving effective noise reduction and improved user experience for open-back acoustic devices.

CN115240697BActive Publication Date: 2026-02-03SHENZHEN SHOKZ CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110486203.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-25
Filing Date
2021-04-30
Publication Date
2026-02-03
Estimated Expiration
2041-08-09

AI Technical Summary

Technical Problem

Existing in-ear acoustic devices block users' ears, causing discomfort, while open-ear acoustic devices have poor noise reduction performance in noisy environments, affecting users' auditory experience.

Method used

A microphone array is used to pick up ambient noise, a processor performs sound field estimation to generate a noise reduction signal, and a target signal is output by a speaker to cancel out the ambient noise. The microphone array is set in the target area to reduce speaker interference, and the speaker can be bone conduction or air conduction type.

Benefits of technology

It effectively reduces ambient noise without blocking the ear canal, improving the user's auditory experience and call quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240697B_ABST
    Figure CN115240697B_ABST
Patent Text Reader

Abstract

An acoustic device is disclosed. The acoustic device can include a microphone array, a processor, and at least one speaker. The microphone array can be configured to pick up ambient noise. The processor can be configured to estimate a sound field of a target spatial location with the microphone array. The target spatial location can be closer to a user's ear canal than any microphone in the microphone array. The processor can be further configured to generate a noise reduction signal based on the picked up ambient noise and the sound field estimate of the target spatial location. The at least one speaker can be configured to output a target signal according to the noise reduction signal. The target signal can be used to reduce the ambient noise. The microphone array can be disposed in a target area to minimize the microphone array from an interfering signal from the at least one speaker.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing

[0002] This application claims priority to international application No. PCT / CN2021 / 089670, filed on April 25, 2021, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of acoustics, and in particular to an acoustic device. Background Technology

[0004] Acoustic devices allow users to listen to audio content and make voice calls while ensuring the privacy of their interactions and without disturbing those around them. Acoustic devices are generally divided into two categories: in-ear and open-ear. In-ear devices block the user's ear during use, and users may experience discomfort such as blockage, foreign body sensation, and pain after prolonged wear. Open-ear devices allow the ear to be open, which is beneficial for long-term wear, but their noise reduction effect is less noticeable in noisy environments, thus reducing the user's auditory experience.

[0005] Therefore, it is desirable to provide an acoustic device that can open up the user's ears and improve the user's auditory experience. Summary of the Invention

[0006] One embodiment of this application provides an acoustic device. The acoustic device may include a microphone array, a processor, and at least one speaker. The microphone array may be configured to pick up ambient noise. The processor may be configured to estimate the sound field of a target spatial location using the microphone array. The target spatial location may be closer to the user's ear canal than any of the microphones in the microphone array. The processor may be further configured to generate a noise-reduced signal based on the picked-up ambient noise and the sound field estimation of the target spatial location. The at least one speaker may be configured to output a target signal based on the noise-reduced signal. The target signal may be used to reduce the ambient noise. The microphone array may be positioned in a target area to minimize interference signals from the at least one speaker.

[0007] In some embodiments, generating a denoised signal based on the sound field estimation of the picked-up ambient noise and the target spatial location may include estimating the noise of the target spatial location based on the picked-up ambient noise and generating the denoised signal based on the noise of the target spatial location and the sound field estimation of the target spatial location.

[0008] In some embodiments, the acoustic device may further include one or more sensors for acquiring motion information of the acoustic device. The processor may be further configured to update the noise and sound field estimate of the target spatial location based on the motion information, and to generate the noise-reduced signal based on the updated noise and sound field estimate of the target spatial location.

[0009] In some embodiments, estimating the target spatial location based on the picked-up environmental noise may include identifying one or more spatial noise sources associated with the picked-up environmental noise and estimating the noise at the target spatial location based on the spatial noise sources.

[0010] In some embodiments, estimating the sound field of a target spatial location using the microphone array may include constructing a virtual microphone based on the microphone array. The virtual microphone includes a mathematical model or machine learning model for representing the audio data collected by the microphone if the target spatial location includes a microphone, and estimating the sound field of the target spatial location based on the virtual microphone.

[0011] In some embodiments, generating a noise-reduced signal based on the picked-up ambient noise and the sound field estimation of the target spatial location may include estimating the noise of the target spatial location based on the virtual microphone and generating the noise-reduced signal based on the noise of the target spatial location and the sound field estimation of the target spatial location.

[0012] In some embodiments, the at least one speaker may be a bone conduction speaker. The interference signal may include leakage sound and vibration signals from the bone conduction speaker. The target region may be the region where the total energy of the leakage sound and vibration signals transmitted to the bone conduction speaker of the microphone array is minimized.

[0013] In some embodiments, the location of the target region may be related to the orientation of the diaphragm of the microphone in the microphone array. The orientation of the microphone diaphragm can reduce the magnitude of the vibration signal from the bone conduction speaker received by the microphone. The orientation of the microphone diaphragm can such that the vibration signal from the bone conduction speaker received by the microphone at least partially cancels out the sound leakage signal from the bone conduction speaker received by the microphone. The vibration signal from the bone conduction speaker received by the microphone can reduce the sound leakage signal from the bone conduction speaker received by the microphone by 5-6 dB.

[0014] In some embodiments, the at least one loudspeaker may be an air-conducting loudspeaker. The target region may be the region with the lowest sound pressure level in the radiated sound field of the air-conducting loudspeaker.

[0015] In some embodiments, the processor may be further configured to process the noise-reduced signal based on a transfer function. The transfer function may include a first transfer function and a second transfer function. The first transfer function may represent the change in parameters of the target signal from the at least one loudspeaker to the location where the target signal and the ambient noise are canceled. The second transfer function may represent the change in parameters of the ambient noise from the target spatial location to the location where the target signal and the ambient noise are canceled. The at least one loudspeaker may be further configured to output the target signal based on the processed noise-reduced signal.

[0016] In some embodiments, generating a noise-reduced signal based on the sound field estimation of the picked-up ambient noise and the target spatial location may include dividing the picked-up ambient noise into multiple frequency bands, the multiple frequency bands corresponding to different frequency ranges, and generating a noise-reduced signal corresponding to each of the at least one frequency band for at least one of the multiple frequency bands.

[0017] In some embodiments, the processor may be further configured to adjust the amplitude and phase of the noise at the target spatial location based on the sound field estimation of the target spatial location to generate the noise-reduced signal.

[0018] In some embodiments, the acoustic device may further include a fixing structure configured to fix the acoustic device in a position near the user's ear without obstructing the user's ear canal.

[0019] In some embodiments, the acoustic device may further include a housing structure configured to carry or house the microphone array, the processor, and the at least one speaker.

[0020] One embodiment of this application provides a noise reduction method. The noise reduction method may include picking up ambient noise using a microphone array. The noise reduction method may include estimating the sound field of a target spatial location using the microphone array by a processor. The target spatial location may be closer to the user's ear canal than any microphone in the microphone array. The noise reduction method may include generating a noise-reduced signal based on the picked-up ambient noise and the sound field estimation of the target spatial location. The noise reduction method may further include outputting a target signal from at least one speaker based on the noise-reduced signal. The target signal can be used to reduce the ambient noise. The microphone array may be positioned in the target area to minimize interference signals from the at least one speaker.

[0021] Some of the additional features of this application will be described in the following description. These additional features will be apparent to those skilled in the art from the following description and accompanying drawings, or from an understanding of the production or operation of the embodiments. The features of this application can be implemented and obtained through practice or by using various aspects of the methods, tools, and combinations set forth in the following detailed examples. Attached Figure Description

[0022] This application will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0023] Figure 1 These are schematic diagrams of the structure of exemplary acoustic devices according to some embodiments of this application;

[0024] Figure 2 This is a schematic diagram of the structure of an exemplary processor according to some embodiments of this application;

[0025] Figure 3 This is an exemplary noise reduction flowchart of an acoustic device according to some embodiments of this application;

[0026] Figure 4 This is an exemplary noise reduction flowchart of an acoustic device according to some embodiments of this application;

[0027] Figure 5A -D is a schematic diagram of an exemplary arrangement of a microphone array according to some embodiments of this application;

[0028] Figure 6A -B is a schematic diagram illustrating an exemplary arrangement of a microphone array according to some embodiments of this application;

[0029] Figure 7 This is an exemplary flowchart illustrating noise for estimating the spatial location of a target according to some embodiments of this application;

[0030] Figure 8 This is a schematic diagram of noise used to estimate the spatial location of a target, according to some embodiments of this application;

[0031] Figure 9 This is an exemplary flowchart illustrating the sound field and noise for estimating the spatial location of a target according to some embodiments of this application;

[0032] Figure 10 This is a schematic diagram illustrating the construction of a virtual microphone according to some embodiments of this application;

[0033] Figure 11This is a schematic diagram of the three-dimensional sound field leakage signal distribution of a bone conduction loudspeaker at 1000Hz, according to some embodiments of this application.

[0034] Figure 12 This is a schematic diagram of the two-dimensional sound field leakage signal distribution of a bone conduction loudspeaker at 1000Hz, according to some embodiments of this application.

[0035] Figure 13 This is a schematic diagram of the frequency response of the total signal of the vibration signal and the leakage signal of a bone conduction loudspeaker according to some embodiments of this application;

[0036] Figure 14A -B is a schematic diagram of the sound field distribution of an air-conducting loudspeaker according to some embodiments of this application;

[0037] Figure 15 This is an exemplary flowchart illustrating the output of a target signal based on a transfer function according to some embodiments of this application; and

[0038] Figure 16 This is an exemplary flowchart illustrating noise for estimating the spatial location of a target according to some embodiments of this application. Detailed Implementation

[0039] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0040] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0041] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0042] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0043] Open-back acoustic devices (such as open-back headphones) are acoustic devices that allow the user's ears to be open. Open-back devices use a fixed structure (e.g., ear hooks, headbands, temples of glasses) to fix the speaker near the user's ears without blocking the ear canal. When using open-back devices, ambient noise can still be heard, resulting in a poorer auditory experience. For example, in noisy environments (e.g., streets, scenic spots), when playing music using open-back devices, ambient noise will directly enter the user's ear canal, causing significant audible noise that interferes with the music listening experience. Similarly, when making calls while wearing open-back devices, the microphone picks up not only the user's own voice but also ambient noise, leading to a poor call experience.

[0044] To address the aforementioned problems, this application provides an acoustic device. The acoustic device may include a microphone array, a processor, and at least one speaker. The microphone array may be configured to pick up ambient noise. The processor may be configured to estimate the sound field at a target spatial location using the microphone array. The target spatial location may be closer to the user's ear canal than any microphone in the microphone array. It is understood that the microphones in the microphone array may be distributed at different locations near the user's ear canal, using each microphone in the microphone array to estimate the sound field near the user's ear canal location (e.g., the target spatial location). The processor may be further configured to generate a noise-reducing signal based on the picked-up ambient noise and the sound field estimation of the target spatial location. At least one speaker may be configured to output a target signal based on the noise-reducing signal. This target signal can be used to reduce ambient noise. Additionally, the microphone array may be positioned in a target area to minimize interference signals from the at least one speaker. When the at least one speaker is a bone conduction speaker, the interference signal may include leakage and vibration signals from the bone conduction speaker, and the target area may be the area with the minimum total energy of the leakage and vibration signals transmitted to the microphone array from the bone conduction speaker. When at least one loudspeaker is an air-conducting loudspeaker, the target area can be the region with the lowest sound pressure level of the radiated sound field of the air-conducting loudspeaker.

[0045] In the embodiments of this application, the above-described configuration utilizes the target signal output by at least one speaker to reduce the ambient noise in the user's ear canal (e.g., the target spatial location), thereby achieving active noise reduction of the acoustic device and improving the user's auditory experience during the use of the acoustic device.

[0046] Furthermore, in embodiments of this application, the microphone array (also referred to as a feedforward microphone) can simultaneously pick up ambient noise and estimate the sound field at the user's ear canal (e.g., the target spatial location).

[0047] In addition, in the embodiments of this application, the microphone array is set in the target area to reduce or avoid the microphone array picking up interference signals (e.g., target signals) emitted by at least one speaker, thereby ensuring the realization of active noise reduction of the open acoustic device.

[0048] Figure 1 This is a schematic diagram illustrating the structure of an exemplary acoustic device 100 according to some embodiments of this application. In some embodiments, the acoustic device 100 may be an open acoustic device. Figure 1 As shown, the acoustic device 100 may include a microphone array 110, a processor 120, and a speaker 130. In some embodiments, the microphone array 110 may pick up ambient noise and convert the picked-up ambient noise into an electrical signal, which is then transmitted to the processor 120 for processing. The processor 120 may be coupled (e.g., electrically connected) to the microphone array 110 and the speaker 130. The processor 120 may receive the electrical signal transmitted from the microphone array 110 and process it to generate a noise-reduced signal, which is then transmitted to the speaker 130. The speaker 130 may output a target signal based on the noise-reduced signal. This target signal may be used to reduce or cancel ambient noise at the user's ear canal location (e.g., a target spatial location), thereby achieving active noise reduction of the acoustic device 100 and improving the user's auditory experience during use of the acoustic device 100.

[0049] Microphone array 110 can be configured to pick up ambient noise. In some embodiments, ambient noise can refer to a combination of various external sounds in the user's environment. By way of example only, ambient noise can include one or more of traffic noise, industrial noise, construction noise, social noise, etc. Traffic noise can include, but is not limited to, the noise of motor vehicles driving and horn honking. Industrial noise can include, but is not limited to, the noise of factory machinery operating. Construction noise can include, but is not limited to, the noise of machinery digging, drilling, and mixing. Social environmental noise can include, but is not limited to, noise from mass gatherings, entertainment and publicity, crowd noise, and noise from household appliances. In some embodiments, microphone array 110 can be positioned near the user's ear canal to pick up ambient noise transmitted to the user's ear canal and convert the picked-up ambient noise into an electrical signal that is transmitted to processor 120 for processing. In some embodiments, microphone array 110 can be positioned at the user's left ear and / or right ear. For example, microphone array 110 can include a first sub-microphone array and a second sub-microphone array. The first sub-microphone array can be located at the user's left ear, and the second sub-microphone array can be located at the user's right ear. The first and second sub-microphone arrays can be operational simultaneously, or one of them can be operational.

[0050] In some embodiments, ambient noise may include the sound of a user speaking. For example, the microphone array 110 may pick up ambient noise based on the call status of the acoustic device 100. When the acoustic device 100 is not in a call state, the sound of the user speaking may be considered ambient noise, and the microphone array 110 may pick up the user speaking sound along with other ambient noise. When the acoustic device 100 is in a call state, the sound of the user speaking may not be considered ambient noise, and the microphone array 110 may pick up ambient noise other than the user speaking sound. For example, the microphone array 110 may pick up noise emitted from a noise source at a certain distance (e.g., 0.5 meters, 1 meter) from the microphone array 110.

[0051] In some embodiments, the microphone array 110 may include one or more air conduction microphones. For example, when a user listens to music using the acoustic device 100, the air conduction microphone can simultaneously acquire ambient noise and the user's voice, treating the acquired ambient noise and the user's voice together as ambient noise. In some embodiments, the microphone array 110 may also include one or more bone conduction microphones. The bone conduction microphone can directly contact the user's skin, and the vibration signals generated by the user's bones or muscles when speaking can be directly transmitted to the bone conduction microphone, which then converts the vibration signals into electrical signals and transmits the electrical signals to the processor 120 for processing. Alternatively, the bone conduction microphone may not directly contact the human body; the vibration signals generated by the user's bones or muscles when speaking can first be transmitted to the housing structure of the acoustic device 100, and then transmitted from the housing structure to the bone conduction microphone. In some embodiments, when the user is in a call, the processor 120 can treat the sound signals acquired by the air conduction microphone as ambient noise and use this ambient noise for noise reduction, while the sound signals acquired by the bone conduction microphone are transmitted as voice signals to the terminal device, thereby ensuring the call quality during the user's call.

[0052] In some embodiments, the processor 120 can control the on / off states of the bone conduction microphone and the air conduction microphone based on the operating state of the acoustic device 100. The operating state of the acoustic device 100 can refer to the usage state when the user wears the acoustic device 100. As an example only, the operating state of the acoustic device 100 can include, but is not limited to, a call state, a non-call state (e.g., music playback state), a voice message sending state, etc. In some embodiments, when the microphone array 110 picks up ambient noise, the on / off states of the bone conduction microphone and the air conduction microphone in the microphone array 110 can be determined according to the operating state of the acoustic device 100. For example, when the user wears the acoustic device 100 to play music, the on / off state of the bone conduction microphone can be a standby state, and the on / off state of the air conduction microphone can be an active state. As another example, when the user wears the acoustic device 100 to send a voice message, the on / off state of both the bone conduction microphone and the air conduction microphone can be an active state. In some embodiments, the processor 120 can control the on / off states of the microphones (e.g., bone conduction microphone and air conduction microphone) in the microphone array 110 by sending control signals.

[0053] In some embodiments, when the acoustic device 100 is in a non-call state (e.g., music playback state), the processor 120 can control the bone conduction microphone to be in standby mode and the air conduction microphone to be in working mode. In the non-call state, the user's own voice signal can be considered as ambient noise. In this case, the user's own voice signal included in the ambient noise picked up by the air conduction microphone can be left unfiltered, allowing the user's own voice signal, as part of the ambient noise, to cancel out the target signal output by the speaker 130. When the acoustic device 100 is in a call state, the processor 120 can control both the bone conduction microphone and the air conduction microphone to be in working mode. In the call state, the user's own voice signal needs to be preserved. In this case, the processor 120 can send a control signal to control the bone conduction microphone to be in working mode, allowing the bone conduction microphone to pick up the user's voice signal. The processor 120 removes the user's voice signal picked up by the bone conduction microphone from the ambient noise picked up by the air conduction microphone, ensuring that the user's own voice signal does not cancel out the target signal output by the speaker 130, thereby guaranteeing normal call functionality for the user.

[0054] In some embodiments, when the acoustic device 100 is in a call state, if the sound pressure level of the ambient noise exceeds a preset threshold, the processor 120 can control the bone conduction microphone to remain operational. The sound pressure level of the ambient noise reflects its intensity. The preset threshold can be a value pre-stored in the acoustic device 100, such as 50dB, 60dB, or 70dB, or any other arbitrary value. When the sound pressure level of the ambient noise exceeds the preset threshold, the ambient noise will affect the user's call quality. The processor 120 can control the bone conduction microphone to remain operational by sending a control signal. The bone conduction microphone can acquire the vibration signal of the user's facial muscles when speaking, while essentially not picking up external ambient noise. The vibration signal picked up by the bone conduction microphone is then used as the voice signal during the call, thereby ensuring normal communication for the user.

[0055] In some embodiments, when the acoustic device 100 is in a call state, if the sound pressure level of the ambient noise is less than a preset threshold, the processor 120 can control the bone conduction microphone to switch from the working state to the standby state. When the sound pressure level of the ambient noise is less than the preset threshold, the sound pressure level of the ambient noise is relatively small compared to the sound pressure level of the sound signal generated by the user's speech. After the sound signal generated by the user's speech, transmitted through the first acoustic path to a certain position of the user's ear, is partially canceled by the target signal output by the speaker 130, transmitted through the second acoustic path to a certain position of the user's ear, the remaining sound signal generated by the user's speech can still be received by the user's auditory center, which is sufficient to ensure the user's normal call. In this case, the processor 120 can control the bone conduction microphone to switch from the working state to the standby state by sending a control signal, thereby reducing the signal processing complexity and the power consumption of the acoustic device 100.

[0056] In some embodiments, depending on the microphone's operating principle, the microphone array 110 may include a dynamic microphone, a ribbon microphone, a condenser microphone, an electret microphone, an electromagnetic microphone, a carbon microphone, or any combination thereof. In some embodiments, the arrangement of the microphone array 110 may include a linear array (e.g., straight line, curved line), a planar array (e.g., regular and / or irregular shapes such as cross, circle, ring, polygon, mesh, etc.), a three-dimensional array (e.g., cylindrical, spherical, hemispherical, polyhedral, etc.), or any combination thereof. Further details regarding the arrangement of the microphone array 110 can be found elsewhere in this application, for example... Figure 5A -D、 Figure 6A -B and its corresponding description.

[0057] Processor 120 can be configured to estimate the sound field of a target spatial location using microphone array 110. The sound field of the target spatial location can refer to the distribution and variation of sound waves at or near the target spatial location (e.g., variation over time, variation over location). Physical quantities describing the sound field can include sound pressure, sound frequency, sound amplitude, sound phase, sound source vibration velocity, or medium (e.g., air) density, etc. Typically, these physical quantities can be functions of position and time. The target spatial location can refer to a spatial location close to a specific distance from the user's ear canal. This target spatial location can be closer to the user's ear canal than any one of the microphones in microphone array 110. This specific distance can be a fixed distance, such as 0.5 cm, 1 cm, 2 cm, 3 cm, etc. In some embodiments, the target spatial location can be related to the number of microphones in microphone array 110 and their distribution relative to the user's ear canal. The target spatial location can be adjusted by adjusting the number of microphones in microphone array 110 and / or their distribution relative to the user's ear canal. For example, increasing the number of microphones in microphone array 110 can make the target spatial location closer to the user's ear canal. For example, the target spatial position can be made closer to the user's ear canal by reducing the spacing between the microphones in the microphone array 110. Alternatively, the target spatial position can be made closer to the user's ear canal by changing the arrangement of the microphones in the microphone array 110.

[0058] Processor 120 may be further configured to generate a denoised signal based on the picked-up ambient noise and a sound field estimation of the target spatial location. Specifically, processor 120 may receive and process the converted electrical signal of ambient noise transmitted from microphone array 110 to obtain parameters of the ambient noise (e.g., amplitude, phase, etc.). Processor 120 may further adjust the parameters of the ambient noise (e.g., amplitude, phase, etc.) based on the sound field estimation of the target spatial location to generate the denoised signal. The parameters of the denoised signal (e.g., amplitude, phase, etc.) correspond to the parameters of the ambient noise. By way of example only, the amplitude of the denoised signal may be approximately equal to the amplitude of the ambient noise, and the phase of the denoised signal may be approximately opposite to the phase of the ambient noise. In some embodiments, processor 120 may include hardware modules and software modules. By way of example only, the hardware module may include a Digital Signal Processor (DSP) chip, an Advanced Reduced Instruction Set Machine (ARM), and the software module may include an algorithm module. Further details about processor 120 can be found elsewhere in this application, for example, Figure 2 And its corresponding description.

[0059] The speaker 130 can be configured to output a target signal based on a noise reduction signal. This target signal can be used to reduce or eliminate ambient noise transmitted to a location on the user's ear (e.g., tympanic membrane, basilar membrane). In some embodiments, when the user wears the acoustic device 100, the speaker 130 can be located near the user's ear. In some embodiments, depending on the speaker's operating principle, the speaker 130 can include one or more of the following: an electrodynamic speaker (e.g., a moving-coil speaker), a magnetic speaker, an ionic speaker, an electrostatic speaker (or a capacitive speaker), a piezoelectric speaker, etc. In some embodiments, depending on the propagation mode of the sound output by the speaker, the speaker 130 can include an air-conduction speaker and / or a bone-conduction speaker. In some embodiments, the number of speakers 130 can be one or more. When there is only one speaker 130, it can be used to output a target signal to eliminate ambient noise and to deliver sound information that the user needs to hear (e.g., device media audio, far-end audio of a call). For example, when there is only one speaker 130 and it is an air-conduction speaker, the air-conduction speaker can be used to output a target signal to eliminate ambient noise. In this scenario, the target signal can be a sound wave (i.e., air vibration), which can be transmitted through the air to the target spatial location and cancel out ambient noise at that location. Simultaneously, the air-conducting speaker can also be used to deliver the sound information the user needs to hear. For example, when there is only one speaker 130 and it is a bone conduction speaker, the bone conduction speaker can be used to output a target signal to eliminate ambient noise. In this case, the target signal can be a vibration signal (e.g., vibration of the speaker housing), which can be transmitted through bone or tissue to the user's basilar membrane and cancel out ambient noise at the user's basilar membrane. Simultaneously, the bone conduction speaker can also be used to deliver the sound information the user needs to hear. When there are multiple speakers 130, some of the multiple speakers 130 can be used to output a target signal to eliminate ambient noise, while others can be used to deliver the sound information the user needs to hear (e.g., device media audio, far-end audio of a call). For example, when there are multiple speakers 130 and they include both bone conduction and air conduction speakers, the air conduction speaker can be used to output sound waves to reduce or eliminate ambient noise, and the bone conduction speaker can be used to deliver the sound information the user needs to hear. Compared to air conduction speakers, bone conduction speakers can transmit mechanical vibrations directly to the user's auditory nerve through the user's body (e.g., bones, skin tissue, etc.), with less interference to the air conduction microphone that picks up ambient noise in the process.

[0060] It should be noted that the speaker 130 can be an independent functional device or part of a single device capable of performing multiple functions. As an example only, the speaker 130 can be integrated with and / or formed as a single unit with the processor 120. In some embodiments, when there are multiple speakers 130, the arrangement of the multiple speakers 130 can include linear arrays (e.g., straight lines, curved lines), planar arrays (e.g., cross-shaped, mesh-shaped, circular, annular, polygonal, etc., regular and / or irregular shapes), three-dimensional arrays (e.g., cylindrical, spherical, hemispherical, polyhedral, etc.), or any combination thereof, which are not limited herein. In some embodiments, the speaker 130 can be positioned at the user's left ear and / or right ear. For example, the speaker 130 can include a first sub-speaker and a second sub-speaker. The first sub-speaker can be located at the user's left ear, and the second sub-speaker can be located at the user's right ear. The first sub-speaker and the second sub-speaker can be in working condition simultaneously, or one of them can be in working condition. In some embodiments, the speaker 130 can be a speaker with a directional sound field, with its main lobe pointing towards the user's ear canal.

[0061] In some embodiments, the acoustic device 100 may further include one or more sensors 140. The one or more sensors 140 may be electrically connected to other components of the acoustic device 100 (e.g., processor 120). The one or more sensors 140 may be used to acquire physical position and / or motion information of the acoustic device 100. By way of example only, the one or more sensors 140 may include an inertial measurement unit (IMU), a global positioning system (GPS), radar, etc. Motion information may include motion trajectory, motion direction, motion speed, motion acceleration, motion angular velocity, motion-related time information (e.g., motion start time, motion end time), etc., or any combination thereof. Taking an IMU as an example, the IMU may include a microelectromechanical system (MEMS). The microelectromechanical system may include a multi-axis accelerometer, a gyroscope, a magnetometer, etc., or any combination thereof. The IMU may be used to detect the physical position and / or motion information of the acoustic device 100 to enable control of the acoustic device 100 based on the physical position and / or motion information. Further details regarding the control of the acoustic device 100 based on physical position and / or motion information can be found elsewhere in this application, for example, Figure 4 And its corresponding description.

[0062] In some embodiments, the acoustic device 100 may include a transceiver 150. The transceiver 150 may be electrically connected to other components of the acoustic device 100 (e.g., processor 120). In some embodiments, the transceiver 150 may include Bluetooth, an antenna, etc. The acoustic device 100 can communicate with other external devices (e.g., mobile phones, tablets, smartwatches) via the transceiver 150. For example, the acoustic device 100 can wirelessly communicate with other devices via Bluetooth.

[0063] In some embodiments, the acoustic device 100 may include a housing structure 160. The housing structure 160 may be configured to carry other components of the acoustic device 100 (e.g., microphone array 110, processor 120, speaker 130, one or more sensors 140, transceiver 150). In some embodiments, the housing structure 160 may be a hollow, closed or semi-closed structure, with other components of the acoustic device 100 located within or on the housing structure. In some embodiments, the housing structure may be a regular or irregular three-dimensional structure, such as a cuboid, cylinder, or frustum. When a user wears the acoustic device 100, the housing structure may be located near the user's ear. For example, the housing structure may be located on the periphery of the user's auricle (e.g., the front or back). Alternatively, the housing structure may be located on the user's ear but not block or cover the user's ear canal. In some embodiments, the acoustic device 100 may be a bone conduction headset, with at least one side of the housing structure in contact with the user's skin. In bone conduction headphones, an acoustic driver (e.g., a vibrating speaker) converts audio signals into mechanical vibrations, which are transmitted to the user's auditory nerve via the housing structure and the user's bones. In some embodiments, the acoustic device 100 may be an air conduction headphone, with at least one side of the housing structure either in contact with or not in contact with the user's skin. At least one sound-guiding hole is included on the sidewall of the housing structure, through which a speaker in the air conduction headphone converts audio signals into air-conducted sound, which is radiated toward the user's ear.

[0064] In some embodiments, the acoustic device 100 may include a fixing structure 170. The fixing structure 170 may be configured to secure the acoustic device 100 near a user's ear without obstructing the user's ear canal. In some embodiments, the fixing structure 170 may be physically connected to the housing structure 160 of the acoustic device 100 (e.g., snap-fit, threaded connection, etc.). In some embodiments, the housing structure 160 of the acoustic device 100 may be part of the fixing structure 170. In some embodiments, the fixing structure 170 may include ear hooks, back hooks, elastic bands, temples, etc., to better secure the acoustic device 100 near the user's ear and prevent it from falling off during use. For example, the fixing structure 170 may be an ear hook, which may be configured to be worn around the ear area. In some embodiments, the ear hook may be a continuous hook and may be elastically stretched to be worn on the user's ear, while also applying pressure to the user's auricle to secure the acoustic device 100 firmly in a specific location on the user's ear or head. In some embodiments, the ear hook may be a discontinuous band. For example, the ear hook may include a rigid portion and a flexible portion. The rigid portion can be made of a rigid material (e.g., plastic or metal) and can be fixed to the housing structure 160 of the acoustic device 100 by a physical connection (e.g., snap-fit, threaded connection, etc.). The flexible portion can be made of an elastic material (e.g., fabric, composite material, and / or neoprene). For example, the fixing structure 170 can be a neck strap configured to be worn around the neck / shoulder area. As another example, the fixing structure 170 can be an eyeglass temple, which, as part of the eyeglasses, is mounted on the user's ear.

[0065] In some embodiments, the acoustic device 100 may further include an interactive module (not shown) for adjusting the sound pressure level of the target signal. In some embodiments, the interactive module may include a button, a voice assistant, a gesture sensor, etc. The user can adjust the noise reduction mode of the acoustic device 100 by controlling the interactive module. Specifically, the user can adjust (e.g., amplify or reduce) the amplitude information of the noise reduction signal by controlling the interactive module to change the sound pressure level of the target signal emitted by the speaker array 130, thereby achieving different noise reduction effects. As an example only, the noise reduction mode may include a strong noise reduction mode, a medium noise reduction mode, a weak noise reduction mode, etc. For example, when the user wears the acoustic device 100 indoors, where the ambient noise is relatively low, the user can turn off the noise reduction mode of the acoustic device 100 or adjust it to a weak noise reduction mode through the interactive module. For example, when a user wears the acoustic device 100 while walking in public places such as on the street, the user needs to maintain a certain level of awareness of the surrounding environment while listening to audio signals (e.g., music, voice information) to cope with emergencies. In this case, the user can select a medium noise reduction mode through an interactive module (e.g., a button or voice assistant) to preserve ambient noise (such as alarm sounds, impact sounds, car horns, etc.). As another example, when a user is taking public transportation such as a subway or airplane, the user can select a strong noise reduction mode through the interactive module to further reduce ambient noise. In some embodiments, the processor 120 can also send prompts to the acoustic device 100 or a terminal device (e.g., a mobile phone, smartwatch, etc.) communicatively connected to the acoustic device 100 based on the ambient noise intensity range to remind the user to adjust the noise reduction mode.

[0066] It should be noted that the above regarding Figure 1 The description provided is for illustrative purposes only and is not intended to limit the scope of this application. Various changes and modifications can be made by those skilled in the art based on the guidance of this application. In some embodiments, one or more components of the acoustic device 100 (e.g., one or more sensors 140, transceiver 150, fixing structure 170, interaction module, etc.) may be omitted. In some embodiments, one or more components of the acoustic device 100 may be replaced by other elements that perform similar functions. For example, the acoustic device 100 may not include the fixing structure 170, and the housing structure 160 or a portion thereof may be a housing structure with a shape adapted to the human ear (e.g., annular, elliptical, polygonal (regular or irregular), U-shaped, V-shaped, semi-circular) so that the housing structure can be attached near the user's ear. In some embodiments, a component of the acoustic device 100 may be divided into multiple sub-components, or multiple components may be combined into a single component. These changes and modifications do not depart from the scope of this application.

[0067] Figure 2This is a schematic diagram of the structure of an exemplary processor 120 according to some embodiments of this application. Figure 2 As shown, the processor 120 may include an analog-to-digital conversion unit 210, a noise estimation unit 220, an amplitude and phase compensation unit 230, and a digital-to-analog conversion unit 240.

[0068] In some embodiments, the analog-to-digital converter (ADC) unit 210 can be configured to convert the signal input from the microphone array 110 into a digital signal. Specifically, the microphone array 110 picks up ambient noise and converts the picked-up ambient noise into an electrical signal, which is then transmitted to the processor 120. Upon receiving the electrical signal of the ambient noise transmitted by the microphone array 110, the ADC unit 210 can convert the electrical signal into a digital signal. In some embodiments, the ADC unit 210 can be electrically connected to the microphone array 110 and further electrically connected to other components of the processor 120 (e.g., the noise estimation unit 220). Furthermore, the ADC unit 210 can transmit the converted digital signal of the ambient noise to the noise estimation unit 220.

[0069] In some embodiments, the noise estimation unit 220 may be configured to estimate ambient noise based on a received digital signal of ambient noise. For example, the noise estimation unit 220 may estimate relevant parameters of the ambient noise at a target spatial location based on the received digital signal of ambient noise. As an example only, these parameters may include the noise source (e.g., location, orientation of the noise source), propagation direction, amplitude, phase, etc., or any combination thereof, of the noise at the target spatial location. In some embodiments, the noise estimation unit 220 may also be configured to estimate the sound field at the target spatial location using the microphone array 110. Further details regarding the estimation of the sound field at the target spatial location can be found elsewhere in this application, for example... Figure 4 And its corresponding description. In some embodiments, the noise estimation unit 220 may be electrically connected to other components of the processor 120 (e.g., the amplitude-phase compensation unit 230). Further, the noise estimation unit 220 may transmit the estimated environmental noise-related parameters and the sound field of the target spatial location to the amplitude-phase compensation unit 230.

[0070] In some embodiments, the amplitude-phase compensation unit 230 can be configured to compensate for parameters related to the estimated environmental noise based on the sound field at the target spatial location. For example, the amplitude-phase compensation unit 230 can compensate for the amplitude and phase of the environmental noise based on the sound field at the target spatial location to obtain a digital noise-reduced signal. In some embodiments, the amplitude-phase compensation unit 230 can adjust the amplitude of the environmental noise and perform inverse compensation for the phase of the environmental noise to obtain a digital noise-reduced signal. The amplitude of the digital noise-reduced signal can be approximately equal to the amplitude of the digital signal corresponding to the environmental noise, and the phase of the digital noise-reduced signal can be approximately opposite to the phase of the digital signal corresponding to the environmental noise. In some embodiments, the amplitude-phase compensation unit 230 can be electrically connected to other components of the processor 120 (e.g., the digital-to-analog converter unit 240). Further, the amplitude-phase compensation unit 230 can transmit the digital noise-reduced signal to the digital-to-analog converter unit 240.

[0071] In some embodiments, the digital-to-analog converter 240 may be configured to convert the digital noise-reduced signal into an analog signal to obtain a noise-reduced signal (e.g., an electrical signal). By way of example only, the digital-to-analog converter 240 may include pulse width modulation (PMW). In some embodiments, the digital-to-analog converter 240 may be electrically connected to other components of the processor 120 (e.g., the speaker 130). Furthermore, the digital-to-analog converter 240 may transmit the noise-reduced signal to the speaker 130.

[0072] In some embodiments, the processor 120 may include a signal amplification unit 250. The signal amplification unit 250 may be configured to amplify an input signal. For example, the signal amplification unit 250 may amplify the signal input from the microphone array 110. As an example only, when the acoustic device 100 is in a call state, the signal amplification unit 250 may be used to amplify the user's speech input from the microphone array 110. As another example, the signal amplification unit 250 may amplify the amplitude of ambient noise based on the sound field of a target spatial location. In some embodiments, the signal amplification unit 250 may be electrically connected to other components of the processor 120 (e.g., the microphone array 110, the noise estimation unit 220, and the amplitude-phase compensation unit 230).

[0073] It should be noted that the above regarding Figure 2The description provided is for illustrative purposes only and is not intended to limit the scope of this application. Various changes and modifications can be made by those skilled in the art based on the guidance of this application. In some embodiments, one or more components in processor 120 (e.g., signal amplification unit 250) may be omitted. In some embodiments, a component of processor 120 may be split into multiple sub-components, or multiple components may be combined into a single component. For example, noise estimation unit 220 and amplitude-phase compensation unit 230 may be integrated into a single component to implement the functions of noise estimation unit 220 and amplitude-phase compensation unit 230. These changes and modifications do not depart from the scope of this application.

[0074] Figure 3 This is an exemplary noise reduction flowchart of an acoustic device according to some embodiments of this application. In some embodiments, process 300 may be performed by acoustic device 100. Figure 3 As shown, process 300 may include:

[0075] In step 310, ambient noise is picked up. In some embodiments, this step may be performed by microphone array 110.

[0076] according to Figure 1 As described in the relevant description, environmental noise can refer to a combination of various external sounds in the user's environment (e.g., traffic noise, industrial noise, construction noise, social noise). In some embodiments, the microphone array 110 can be located near the user's ear canal to pick up the environmental noise transmitted to the user's ear canal. Furthermore, the microphone array 110 can convert the picked-up environmental noise signal into an electrical signal and transmit it to the processor 120 for processing.

[0077] In step 320, the noise of the target's spatial location is estimated based on the picked-up ambient noise. In some embodiments, this step may be performed by processor 120.

[0078] In some embodiments, the processor 120 can perform signal separation on the picked-up ambient noise. In some embodiments, the ambient noise picked up by the microphone array 110 may include various sounds. The processor 120 can perform signal analysis on the ambient noise picked up by the microphone array 110 to separate the various sounds. Specifically, the processor 120 can adaptively adjust the parameters of the filter based on the statistical distribution characteristics and structural features of various sounds in different dimensions such as space, time domain, and frequency domain, estimate the parameter information of each sound signal in the ambient noise, and complete the signal separation process based on the parameter information of each sound signal. In some embodiments, the statistical distribution characteristics of noise may include probability distribution density, power spectral density, autocorrelation function, probability density function, variance, mathematical expectation, etc. In some embodiments, the structural features of noise may include noise distribution, noise intensity, global noise intensity, noise rate, etc., or any combination thereof. Global noise intensity may refer to average noise intensity or weighted average noise intensity. Noise rate may refer to the degree of dispersion of noise distribution. As an example only, the ambient noise picked up by the microphone array 110 may include a first signal, a second signal, and a third signal. Processor 120 acquires the differences between the first signal, the second signal, and the third signal in the spatial domain (e.g., signal location), the temporal domain (e.g., delay), and the frequency domain (e.g., amplitude, phase), and separates the first signal, the second signal, and the third signal according to the differences in these three dimensions, obtaining relatively pure first signal, second signal, and third signal. Further, processor 120 can update the ambient noise based on the parameter information (e.g., frequency information, phase information, amplitude information) of the separated signals. For example, processor 120 can determine that the first signal is the user's call voice based on the parameter information of the first signal, and remove the first signal from the ambient noise to update the ambient noise. In some embodiments, the removed first signal can be transmitted to the far end of the call. For example, when a user wears the acoustic device 100 to make a voice call, the first signal can be transmitted to the far end of the call.

[0079] The target spatial location is a location in or near the user's ear canal, determined based on the microphone array 110. Figure 1 As described in the relevant description, the target spatial location can refer to a spatial location within a specific distance (e.g., 0.5cm, 1cm, 2cm, 3cm) of the user's ear canal (e.g., ear opening). In some embodiments, the target spatial location is closer to the user's ear canal than any of the microphones in the microphone array 110. Figure 1As described in the relevant description, the target spatial location is related to the number of microphones in the microphone array 110 and their distribution relative to the user's ear canal. The target spatial location can be adjusted by adjusting the number of microphones in the microphone array 110 and / or their distribution relative to the user's ear canal. In some embodiments, estimating the noise of the target spatial location based on the picked-up ambient noise (or updated ambient noise) may further include identifying one or more spatial noise sources related to the picked-up ambient noise and estimating the noise of the target spatial location based on the spatial noise sources. The ambient noise picked up by the microphone array 110 may come from spatial noise sources of different orientations and types. The parameter information (e.g., frequency information, phase information, amplitude information) corresponding to each spatial noise source is different. In some embodiments, the processor 120 can perform signal separation and extraction of the noise at the target spatial location according to the statistical distribution and structural characteristics of different types of noise in different dimensions (e.g., spatial domain, time domain, frequency domain, etc.), thereby obtaining different types of noise (e.g., different frequencies, different phases, etc.) and estimating the parameter information (e.g., amplitude information, phase information, etc.) corresponding to each type of noise. In some embodiments, the processor 120 may further determine the overall parameter information of the noise at the target spatial location based on the parameter information corresponding to different types of noise at the target spatial location. More information regarding estimating the noise at the target spatial location based on one or more spatial noise sources can be found elsewhere in this specification, for example... Figure 7-8 And its corresponding description.

[0080] In some embodiments, estimating the target spatial location based on the picked-up ambient noise (or updated ambient noise) may further include constructing a virtual microphone based on the microphone array 110 and estimating the target spatial location based on the virtual microphone. More details regarding estimating the target spatial location based on the virtual microphone can be found elsewhere in this specification, such as... Figure 9-10 And its corresponding description.

[0081] In step 330, a noise-reduced signal is generated based on the noise at the target spatial location. In some embodiments, this step may be performed by processor 120.

[0082] In some embodiments, the processor 120 can generate a noise-reduced signal based on the parameter information (e.g., amplitude information, phase information, etc.) of the noise at the target spatial location obtained in step 320. In some embodiments, the phase difference between the phase of the noise-reduced signal and the phase of the noise at the target spatial location can be less than or equal to a preset phase threshold. The preset phase threshold can be in the range of 90-180 degrees. The preset phase threshold can be adjusted within this range according to the user's needs. For example, when the user does not want to be disturbed by ambient noise, the preset phase threshold can be a larger value, such as 180 degrees, that is, the phase of the noise-reduced signal is opposite to the phase of the noise at the target spatial location. As another example, when the user wants to remain sensitive to the ambient noise, the preset phase threshold can be a smaller value, such as 90 degrees. It should be noted that the more ambient noise the user wants to receive, the closer the preset phase threshold can be to 90 degrees, and the less ambient noise the user wants to receive, the closer the preset phase threshold can be to 180 degrees. In some embodiments, when the phase of the noise-reduced signal is constant with the phase of the noise at the target spatial location (e.g., out of phase), the amplitude difference between the amplitude of the noise at the target spatial location and the amplitude of the noise-reduced signal can be less than or equal to a preset amplitude threshold. For example, when a user does not want to be disturbed by ambient noise, the preset amplitude threshold can be a small value, such as 0 dB, meaning the amplitude of the noise-reduced signal is equal to the amplitude of the noise at the target spatial location. Alternatively, when a user wants to remain sensitive to their surroundings, the preset amplitude threshold can be a large value, such as approximately equal to the amplitude of the noise at the target spatial location. It should be noted that the more ambient noise a user wants to receive, the closer the preset amplitude threshold can be to the amplitude of the noise at the target spatial location; conversely, the less ambient noise a user wants to receive, the closer the preset amplitude threshold can be to 0 dB.

[0083] In some embodiments, the loudspeaker 130 can output a target signal based on a noise-reducing signal generated by the processor 120. For example, the loudspeaker 130 can convert the noise-reducing signal (e.g., an electrical signal) into a target signal (i.e., a vibration signal) based on a vibrating component in the loudspeaker 130, which can cancel out ambient noise. In some embodiments, when the noise at the target spatial location consists of multiple spatial noise sources, the loudspeaker 130 can output a target signal corresponding to each of the multiple spatial noise sources based on the noise-reducing signal. For example, if the multiple spatial noise sources include a first spatial noise source and a second spatial noise source, the loudspeaker 130 can output a first target signal with approximately opposite phase and approximately equal amplitude to the noise of the first spatial noise source to cancel out the noise of the first spatial noise source, and a second target signal with approximately opposite phase and approximately equal amplitude to the noise of the second spatial noise source to cancel out the noise of the second spatial noise source. In some embodiments, when the loudspeaker 130 is an air-conducting loudspeaker, the location where the target signal cancels out ambient noise can be the target spatial location. The distance between the target spatial location and the user's ear canal is small, and the noise at the target spatial location can be approximated as the noise at the user's ear canal location. Therefore, the noise reduction signal and the noise at the target spatial location cancel each other out, which can be approximated as the elimination of environmental noise transmitted to the user's ear canal, thus achieving active noise reduction of the acoustic device 100. In some embodiments, when the speaker 130 is a bone conduction speaker, the location where the target signal and environmental noise cancel each other out can be the basilar membrane. The target signal and environmental noise are canceled out at the user's basilar membrane, thereby achieving active noise reduction of the acoustic device 100.

[0084] It should be noted that the above description of process 300 is merely for illustration and explanation, and does not limit the scope of this application. Those skilled in the art can make various modifications and changes to process 300 under the guidance of this application. For example, steps in process 300 can be added, omitted, or merged. Furthermore, signal processing (e.g., filtering) can be performed on environmental noise. These modifications and changes are still within the scope of this application.

[0085] Figure 4 This is an exemplary noise reduction flowchart of an acoustic device according to some embodiments of this application. In some embodiments, process 400 may be performed by acoustic device 100. Figure 4 As shown, process 400 may include:

[0086] In step 410, ambient noise is picked up. In some embodiments, this step may be performed by microphone array 110. In some embodiments, step 410 may be performed in a similar manner to step 310, and the relevant description will not be repeated here.

[0087] In step 420, the noise of the target's spatial location is estimated based on the picked-up ambient noise. In some embodiments, this step may be performed by processor 120. In some embodiments, step 420 may be performed in a similar manner to step 320, and the relevant description will not be repeated here.

[0088] In step 430, the sound field at the target spatial location is estimated. In some embodiments, this step may be performed by processor 120.

[0089] In some embodiments, the processor 120 may utilize the microphone array 110 to estimate the sound field at the target spatial location. Specifically, the processor 120 may construct a virtual microphone based on the microphone array 110 and estimate the sound field at the target spatial location based on the virtual microphone. Further details regarding the estimation of the sound field at the target spatial location based on a virtual microphone can be found elsewhere in this specification, for example... Figure 9-10 And its corresponding description.

[0090] In step 440, a denoised signal is generated based on the noise at the target spatial location and the sound field estimation at the target spatial location. In some embodiments, step 440 may be executed by processor 120.

[0091] In some embodiments, the processor 120 can adjust the noise parameter information (e.g., frequency information, amplitude information, phase information) of the target spatial location based on the physical quantities (e.g., sound pressure, sound frequency, sound amplitude, sound phase, sound source vibration velocity, or medium (e.g., air) density) related to the sound field at the target spatial location obtained in step 430 to generate a noise-reduced signal. For example, the processor 120 can determine whether the physical quantities (e.g., sound frequency, sound amplitude, sound phase) related to the sound field are the same as the noise parameter information of the target spatial location. If the physical quantities related to the sound field are the same as the noise parameter information of the target spatial location, the processor 120 may not adjust the noise parameter information of the target spatial location. If the physical quantities related to the sound field are different from the noise parameter information of the target spatial location, the processor 120 can determine the difference between the physical quantities related to the sound field and the noise parameter information of the target spatial location, and adjust the noise parameter information of the target spatial location based on the difference. As an example only, when the difference exceeds a certain range, the processor 120 can use the average of the physical quantities related to the sound field and the parameter information of the noise at the target spatial location as the adjusted parameter information of the noise at the target spatial location, and generate a noise-reduced signal based on the adjusted parameter information of the noise at the target spatial location. For another example, since noise in the environment is constantly changing, when the processor 120 generates the noise-reduced signal, the noise at the target spatial location in the actual environment may have undergone subtle changes. Therefore, the processor 120 can estimate the change in the parameter information of the environmental noise at the target spatial location based on the time information of the environmental noise picked up by the microphone array, the current time information, and the physical quantities related to the sound field at the target spatial location (e.g., sound source vibration velocity, medium (e.g., air) density), and adjust the parameter information of the noise at the target spatial location based on this change. The above adjustments make the amplitude and frequency information of the noise reduction signal more consistent with the amplitude and frequency information of the ambient noise at the current target spatial location, and the phase information of the noise reduction signal more consistent with the anti-phase information of the ambient noise at the current target spatial location. This allows the noise reduction signal to eliminate ambient noise more accurately, improving the noise reduction effect and the user's auditory experience.

[0092] In some embodiments, when the position of the acoustic device 100 changes, for example, when the head of a user wearing the acoustic device 100 rotates, the ambient noise (e.g., noise direction, amplitude, phase) changes accordingly. The speed at which the acoustic device 100 performs noise reduction may not keep up with the speed of change in ambient noise, leading to failure of the active noise reduction function or even increased noise. To address this, the processor 120 can update the noise at the target spatial location and the sound field estimate at the target spatial location based on the motion information of the acoustic device 100 (e.g., motion trajectory, motion direction, motion speed, motion acceleration, motion angular velocity, motion-related time information) acquired by one or more sensors 140 of the acoustic device 100. Furthermore, based on the updated noise at the target spatial location and the sound field estimate at the target spatial location, the processor 120 can generate a noise-reduced signal. One or more sensors 140 can record the motion information of the acoustic device 100, allowing the processor 120 to quickly update the noise-reduced signal. This improves the noise tracking performance of the acoustic device 100, enabling the noise-reduced signal to more accurately eliminate ambient noise, further improving the noise reduction effect and the user's auditory experience.

[0093] In some embodiments, the processor 120 can divide the picked-up ambient noise into multiple frequency bands. These multiple frequency bands correspond to different frequency ranges. For example, the processor 120 can divide the picked-up ambient noise into four frequency bands: 100-300Hz, 300-500Hz, 500-800Hz, and 800-1500Hz. In some embodiments, each frequency band includes parameter information (e.g., frequency information, amplitude information, phase information) of the ambient noise within the corresponding frequency range. For at least one of the multiple frequency bands, the processor 120 can perform steps 420-440 to generate a noise-reduced signal corresponding to each of the at least one frequency band. For example, the processor 120 can perform steps 420-440 on frequency bands 300-500Hz and 500-800Hz to generate noise-reduced signals corresponding to frequency bands 300-500Hz and 500-800Hz, respectively. Further, in some embodiments, the speaker 130 can output a target signal corresponding to each frequency band based on the noise-reduced signal corresponding to each frequency band. For example, the loudspeaker 130 can output a target signal that is approximately opposite in phase and approximately equal in amplitude to the noise in the 300-500Hz frequency band to cancel the noise in the 300-500Hz frequency band, and a target signal that is approximately opposite in phase and approximately equal in amplitude to the noise in the 500-800Hz frequency band to cancel the noise in the 500-800Hz frequency band.

[0094] In some embodiments, the processor 120 can also update the noise reduction signal based on manual input from the user. For example, when a user wears the acoustic device 100 to play music in a noisy environment, their auditory experience may be unsatisfactory. The user can manually adjust the parameters of the noise reduction signal (e.g., frequency, phase, and amplitude information) based on their own auditory experience. As another example, when a special user (e.g., a hearing-impaired user or an elderly user) uses the acoustic device 100, their hearing ability differs from that of a normal user. The noise reduction signal generated by the acoustic device 100 itself may not meet the needs of the special user, resulting in a poor auditory experience. In this case, some adjustment factors for the noise reduction signal parameters can be preset. The special user can adjust the noise reduction signal based on their own auditory experience and the preset adjustment factors, thereby updating the noise reduction signal to improve their auditory experience. In some embodiments, the user can manually adjust the noise reduction signal using buttons on the acoustic device 100. In other embodiments, the user can adjust the noise reduction signal through a terminal device. Specifically, the acoustic device 100 or an external device (e.g., a mobile phone, tablet, or computer) that is communicatively connected to the acoustic device 100 can display the parameter information of the suggested noise reduction signal to the user, and the user can fine-tune the parameter information according to their own auditory experience.

[0095] It should be noted that the above description of process 400 is merely for illustration and explanation, and does not limit the scope of this application. Those skilled in the art can make various modifications and changes to process 400 under the guidance of this application. For example, steps in process 400 can be added, omitted, or merged. These modifications and changes are still within the scope of this application.

[0096] Figure 5A -D is a schematic diagram illustrating an exemplary arrangement of a microphone array (e.g., microphone array 110) according to some embodiments of this application. In some embodiments, the microphone array arrangement may be a regular geometric shape. Figure 5A As shown, the microphone array can be a linear array. In some embodiments, the microphone array can also be arranged in other shapes. For example, as... Figure 5B As shown, the microphone array can be a cross-shaped array. For example, such as... Figure 5C As shown, the microphone array can be a circular array. In some embodiments, the microphone array can also be arranged in an irregular geometric shape. For example, as... Figure 5D As shown, the microphone array can be an irregular array. It should be noted that the arrangement of the microphone array is not limited to... Figure 5AThe linear array, cross array, circular array, and irregular array shown in -D can also be arrays of other shapes, such as triangular arrays, spiral arrays, planar arrays, three-dimensional arrays, radial arrays, etc. This application does not limit the types of arrays.

[0097] In some embodiments, Figure 5A Each short solid line in -D can be considered as a microphone or a group of microphones. When each short solid line is considered as a group of microphones, the number of microphones in each group can be the same or different, the type of microphones in each group can be the same or different, and the orientation of each group of microphones can be the same or different. The type, number, and orientation of the microphones can be adapted to the actual application, and this application does not limit them.

[0098] In some embodiments, the microphones in the microphone array may be uniformly distributed. Uniform distribution here can mean that the spacing between any two adjacent microphones in the microphone array is the same. In some embodiments, the microphones in the microphone array may also be non-uniformly distributed. Non-uniform distribution here can mean that the spacing between any two adjacent microphones in the microphone array is different. The spacing between the microphones in the microphone array can be adjusted adaptively according to actual conditions, and this application does not limit this.

[0099] Figure 6A -B is a schematic diagram illustrating an exemplary arrangement of a microphone array (e.g., microphone array 110) according to some embodiments of this application. Figure 6A As shown, when a user wears an acoustic device with a microphone array, the microphone array is arranged in a semi-circular pattern at or around the user's ear, such as... Figure 6B As shown, the microphone array is arranged in a linear pattern at the ear level. It should be noted that the arrangement of the microphone array is not limited to this. Figure 6A and Figure 6B The semi-circular and linear microphone arrays shown are not limited to specific placement positions. Figure 6A and Figure 6B The locations shown, including the semicircles, lines, and microphone array placement, are for illustrative purposes only.

[0100] Figure 7 This is an exemplary flowchart illustrating noise for estimating the spatial location of a target according to some embodiments of this application. Figure 7 As shown, process 700 may include:

[0101] In step 710, one or more spatial noise sources related to the ambient noise picked up by the microphone array are identified. In some embodiments, this step may be performed by processor 120. As described herein, identifying spatial noise sources means determining information related to the spatial noise sources, such as the location of the spatial noise sources (including the orientation of the spatial noise sources, the distance between the spatial noise sources and the target spatial location, etc.), the phase of the spatial noise sources, and the amplitude of the spatial noise sources.

[0102] In some embodiments, spatial noise sources related to environmental noise refer to noise sources whose sound waves can reach or approach the user's ear canal (e.g., the target spatial location). In some embodiments, spatial noise sources can be noise sources in different directions of the user's body (e.g., in front, behind, etc.). For example, if there is noise from a crowd in front of the user and vehicle horn noise to the left of the user, the spatial noise sources include the noise from the crowd in front of the user and the vehicle horn noise to the left of the user. In some embodiments, a microphone array (e.g., microphone array 110) can pick up spatial noise from various directions of the user's body and convert the spatial noise into electrical signals, which are then transmitted to the processor 120. The processor 120 can analyze the electrical signals corresponding to the spatial noise to obtain parameter information (e.g., frequency information, amplitude information, phase information, etc.) of the picked-up spatial noise from each direction. The processor 120 determines the information of the spatial noise sources in each direction based on the parameter information of the spatial noise from each direction, such as the orientation of the spatial noise source, the distance of the spatial noise source, the phase of the spatial noise source, and the amplitude of the spatial noise source. In some embodiments, the processor 120 can determine the spatial noise source based on spatial noise picked up by a microphone array (e.g., microphone array 110) using a noise localization algorithm. The noise localization algorithm may include one or more of beamforming algorithms, super-resolution spatial spectrum estimation algorithms, and time difference of arrival algorithms (also known as time delay estimation algorithms). A beamforming algorithm is a sound source localization method based on controllable beamforming with maximum output power. As an example only, beamforming algorithms may include Steering Response Power-Phase Transform (SPR-PHAT) algorithms, delay-and-sum beamforming, differential microphone algorithms, Generalized Sidelobe Canceller (GSC) algorithms, and Minimum Variance Distortionless Response (MVDR) algorithms. Super-resolution spatial spectrum estimation algorithms can include autoregressive AR models, minimum variance spectrum estimation (MV), and eigenvalue decomposition methods (e.g., multiple signal classification (MUSIC) algorithms). These methods can calculate the correlation matrix of the spatial spectrum by acquiring the sound signal (e.g., spatial noise) picked up by the microphone array and effectively estimate the direction of the spatial noise source.The time difference of arrival algorithm can first estimate the sound arrival time difference and obtain the acoustic delay (TDOA) between microphones in the microphone array. Then, using the obtained sound arrival time difference and the known spatial location of the microphone array, it can further locate the spatial noise source.

[0103] For example, time delay estimation algorithms can calculate the time difference between the arrival of ambient noise signals at different microphones in a microphone array, and then determine the location of the noise source through geometric relationships. Another example is the SPR-PHAT algorithm, which can perform beamforming in the direction of each noise source, and the direction with the strongest beam energy can be approximated as the direction of the noise source. Yet another example is the MUSIC algorithm, which can separate the direction of the ambient noise by performing eigenvalue decomposition on the covariance matrix of the ambient noise signal picked up by the microphone array to obtain a subspace of the ambient noise signal. More information on determining noise sources can be found elsewhere in this application specification, for example... Figure 8 And its corresponding description.

[0104] In some embodiments, a spatial super-resolution image of environmental noise can be formed by methods such as synthetic aperture, sparse recovery, and coprime array. This spatial super-resolution image can be used to reflect the signal reflection map of environmental noise, thereby further improving the positioning accuracy of spatial noise sources.

[0105] In some embodiments, the processor 120 can divide the picked-up ambient noise into multiple frequency bands according to a specific bandwidth (e.g., each band is 500 Hz), each band corresponding to a different frequency range, and determine the spatial noise source corresponding to at least one frequency band. For example, the processor 120 can perform signal analysis on the frequency bands of the ambient noise division to obtain parameter information of the ambient noise corresponding to each frequency band, and determine the spatial noise source corresponding to each frequency band based on the parameter information. As another example, the processor 120 can determine the spatial noise source corresponding to each frequency band using a noise localization algorithm.

[0106] In step 720, noise at the target spatial location is estimated based on the spatial noise source. In some embodiments, this step may be performed by processor 120. As described herein, estimating the noise at the target spatial location refers to estimating parametric information of the noise at the target spatial location, such as frequency information, amplitude information, phase information, etc.

[0107] In some embodiments, the processor 120 can estimate the parameter information of the noise transmitted from each spatial noise source to the target spatial location based on the parameter information (e.g., frequency information, amplitude information, phase information, etc.) of spatial noise sources located in various directions of the user's body obtained in step 710, thereby estimating the noise at the target spatial location. For example, there is a spatial noise source in a first direction (e.g., in front) and a second direction (e.g., behind) of the user's body. The processor 120 can estimate the frequency information, phase information, or amplitude information of the spatial noise source in the first direction when the noise is transmitted to the target spatial location based on the location information, frequency information, phase information, or amplitude information of the spatial noise source in the first direction. The processor 120 can estimate the frequency information, phase information, or amplitude information of the spatial noise source in the second direction when the noise is transmitted to the target spatial location based on the location information, frequency information, phase information, or amplitude information of the spatial noise source in the second direction. Furthermore, the processor 120 can estimate the noise information of the target spatial location based on the frequency, phase, or amplitude information of the first and second directional spatial noise sources, thereby estimating the noise information of the target spatial location. As an example only, the processor 120 can utilize virtual microphone technology or other methods to estimate the noise information of the target spatial location. In some embodiments, the processor 120 can extract the noise parameter information of the spatial noise source from the frequency response curve of the spatial noise source picked up by the microphone array using feature extraction methods. In some embodiments, the methods for extracting the noise parameter information of the spatial noise source may include, but are not limited to, Principal Components Analysis (PCA), Independent Component Algorithm (ICA), Linear Discriminant Analysis (LDA), and Singular Value Decomposition (SVD).

[0108] It should be noted that the above description of process 700 is merely for illustration and explanation, and does not limit the scope of this application. Those skilled in the art can make various modifications and changes to process 700 under the guidance of this application. For example, process 700 may also include steps such as locating the spatial noise source and extracting the noise parameter information of the spatial noise source. As another example, steps 710 and 720 may be combined into one step. These modifications and changes are still within the scope of this application.

[0109] Figure 8This is a schematic diagram illustrating noise for estimating the spatial location of a target, based on some embodiments of this application. The following uses the time difference of arrival (TDOA) algorithm as an example to explain how the localization of spatial noise sources is achieved. Figure 8 As shown, the processor (e.g., processor 120) can calculate the time difference between the noise signal generated by the noise source (e.g., 811, 812, 813) and the different microphones (e.g., microphone 821, microphone 822, etc.) in the microphone array 820, and then, in combination with the known spatial position of the microphone array 820, determine the position of the noise source by the positional relationship (e.g., distance, relative orientation) between the microphone array 820 and the noise source.

[0110] After obtaining the location of the noise sources (e.g., 811, 812, 813), the processor can estimate the phase delay and amplitude change of the noise signal emitted by the noise sources as it travels from the noise sources to the target spatial location 830. Based on this phase delay, amplitude change, and parameter information of the noise signal emitted by the spatial noise sources (e.g., frequency information, amplitude information, phase information, etc.), the processor can obtain parameter information (e.g., frequency information, amplitude information, phase information, etc.) of the environmental noise as it travels to the target spatial location 830, thereby estimating the noise at the target spatial location.

[0111] It should be noted that, Figure 8 The noise sources 811, 812, and 813, the microphone array 820, and the microphones 821 and 822 within the microphone array 820, and the target spatial location 830 described herein are merely illustrative and not intended to limit the scope of this application. Various modifications and alterations can be made by those skilled in the art under the guidance of this application. For example, the microphones in the microphone array 820 are not limited to microphones 821 and 822; the microphone array 820 may also include more microphones, etc. These modifications and alterations are still within the scope of this application.

[0112] Figure 9 This is an exemplary flowchart illustrating the estimation of noise and sound field of a target spatial location according to some embodiments of this application. Figure 9 As shown, process 900 may include:

[0113] In step 910, a virtual microphone is constructed based on a microphone array (e.g., microphone array 110, microphone array 820). In some embodiments, this step may be performed by a processor 120.

[0114] In some embodiments, a virtual microphone can be used to represent or simulate audio data collected by a microphone if a microphone is placed at the target spatial location. That is, the audio data obtained through a virtual microphone can be approximately or equivalent to the audio data collected by a physical microphone if a physical microphone is placed at the target spatial location.

[0115] In some embodiments, the virtual microphone may include a mathematical model. This mathematical model can reflect the relationship between the noise or sound field estimate of the target spatial location and the parameter information (e.g., frequency, amplitude, phase, etc.) of the ambient noise picked up by the microphone array and the parameters of the microphone array. The parameters of the microphone array may include one or more of the following: the arrangement of the microphone array, the spacing between the microphones, the number and position of the microphones in the microphone array, etc. The mathematical model can be calculated based on an initial mathematical model and the parameters of the microphone array and the parameter information (e.g., frequency, amplitude, phase, etc.) of the sound picked up by the microphone array (e.g., ambient noise). For example, the initial mathematical model may include parameters corresponding to the parameters of the microphone array and the parameter information of the ambient noise picked up by the microphone array, as well as model parameters. The initial values ​​of the microphone array parameters, the parameter information of the sound picked up by the microphone array, and the initial values ​​of the model parameters are input into the initial mathematical model to obtain the predicted noise or sound field of the target spatial location. This predicted noise or sound field is then compared with the data (noise and sound field estimates) obtained by a physical microphone positioned at the target spatial location to adjust the model parameters of the mathematical model. Based on the above adjustment method, the mathematical model is obtained by repeatedly adjusting a large amount of data (e.g., the parameters of the microphone array and the parameters of the ambient noise picked up by the microphone array).

[0116] In some embodiments, the virtual microphone may include a machine learning model. This machine learning model can be obtained through training based on the parameters of the microphone array and the parameter information (e.g., frequency, amplitude, phase, etc.) of the sound picked up by the microphone array (e.g., ambient noise). For example, the machine learning model is obtained by training an initial machine learning model (e.g., a neural network model) using the parameters of the microphone array and the parameter information of the sound picked up by the microphone array as training samples. Specifically, the parameters of the microphone array and the parameter information of the sound picked up by the microphone array can be input into the initial machine learning model to obtain prediction results (e.g., noise and sound field estimates at the target spatial location). Then, the prediction results are compared with the data (noise and sound field estimates) obtained by a physical microphone set up at the target spatial location to adjust the parameters of the initial machine learning model. Based on the above adjustment method, through a large amount of data (e.g., the parameters of the microphone array and the parameter information of the ambient noise picked up by the microphone array), and through multiple iterations, the parameters of the initial machine learning model are optimized until the prediction results of the initial machine learning model are the same as or approximately the same as the data obtained by the physical microphone set up at the target spatial location, thus obtaining the machine learning model.

[0117] Virtual microphone technology can move physical microphones away from locations where placement is difficult (e.g., target spatial locations). For example, to ensure open ears without obstructing ear canals, physical microphones cannot be placed at the user's ear canal location (e.g., target spatial location). In this case, virtual microphone technology can be used to place a microphone array close to the user's ears without obstructing ear canals, such as at the auricle, and then construct a virtual microphone positioned at the user's ear canal location using the microphone array. The virtual microphone can use the physical microphone (i.e., the microphone array) at a first location to predict sound data (e.g., amplitude, phase, sound pressure level, sound field, etc.) at a second location (e.g., target spatial location). In some embodiments, the sound data predicted by the virtual microphone at the second location (also referred to as a specific location, such as the target spatial location) can be adjusted based on the distance between the virtual microphone and the physical microphone (i.e., the microphone array), the type of virtual microphone (e.g., mathematical model virtual microphone, machine learning virtual microphone), etc. For example, the closer the virtual microphone is to the physical microphone (i.e., the microphone array), the more accurate the sound data predicted by the virtual microphone at the second location. For example, in certain application scenarios, the sound data of the second position predicted by a machine learning virtual microphone is more accurate than that of a mathematical model virtual microphone. In some embodiments, the position corresponding to the virtual microphone (i.e., the second position, such as the target spatial position) can be near or far from the microphone array.

[0118] In step 920, noise and sound field at the target's spatial location are estimated based on the virtual microphone. In some embodiments, this step may be performed by processor 120.

[0119] In some embodiments, when the virtual microphone is a mathematical model, the processor 120 can input the parameter information of the ambient noise picked up by the microphone array (e.g., frequency information, amplitude information, phase information, etc.) and the parameters of the microphone array (e.g., the arrangement of the microphone array, the spacing between each microphone, and the number of microphones in the microphone array) as parameters of the mathematical model in real time to estimate the noise and sound field of the target spatial location.

[0120] In some embodiments, when the virtual microphone is a machine learning model, the processor 120 can input the environmental noise parameter information (e.g., frequency information, amplitude information, phase information, etc.) picked up by the microphone array and the parameters of the microphone array (e.g., the arrangement of the microphone array, the spacing between each microphone, the number of microphones in the microphone array) into the machine learning model in real time, and estimate the noise and sound field of the target spatial location based on the output of the machine learning model.

[0121] It should be noted that the above description of process 900 is merely for illustration and explanation, and does not limit the scope of this application. Those skilled in the art can make various modifications and changes to process 900 under the guidance of this application. For example, step 920 can be divided into two steps to estimate the noise and sound field of the target spatial location separately. These modifications and changes are still within the scope of this application.

[0122] Figure 10 This is a schematic diagram illustrating the construction of a virtual microphone according to some embodiments of this application. For example... Figure 10 As shown, the target spatial location 1010 can be located near the user's ear canal. In order to achieve the goal of opening the user's ears without blocking the ear canals, a physical microphone cannot be set at the target spatial location 1010, so the noise and sound field of the target spatial location 1010 cannot be directly estimated through a physical microphone.

[0123] To estimate the noise and sound field at target spatial location 1010, a microphone array 1020 can be placed near target spatial location 1010. This is just an example. Figure 10 As shown, the microphone array 1020 may include a first microphone 1021, a second microphone 1022, and a third microphone 1023. Each microphone in the microphone array 1020 (e.g., the first microphone 1021, the second microphone 1022, and the third microphone 1023) can pick up ambient noise in the user's space. Based on the parameter information of the ambient noise picked up by each microphone in the microphone array 1020 (e.g., frequency information, amplitude information, phase information, etc.) and the parameters of the microphone array 1020 (e.g., the arrangement of the microphone array 1020, the spacing between each microphone, and the number of microphones in the microphone array 1020), the processor 120 can construct a virtual microphone. Furthermore, based on this virtual microphone, the processor 120 can estimate the noise and sound field at the target spatial location 1010.

[0124] It should be noted that, Figure 10 The target spatial location 1010 and microphone array 1020, as well as the first microphone 1021, second microphone 1022, and third microphone 1023 in microphone array 1020, described herein are merely examples and illustrations and do not limit the scope of this application. Those skilled in the art can make various modifications and changes under the guidance of this application. For example, the microphones in microphone array 1020 are not limited to the first microphone 1021, second microphone 1022, and third microphone 1023; microphone array 1020 may also include more microphones, etc. These modifications and changes are still within the scope of this application.

[0125] In some embodiments, microphone arrays (e.g., microphone array 110, microphone array 820, microphone array 1020) may pick up interference signals emitted by speakers (e.g., target signals and other sound signals) while picking up ambient noise. To avoid the microphone array picking up interference signals from speakers, the microphone array can be positioned away from the speakers. However, when positioned away from the speakers, the microphone array may be unable to accurately estimate the sound field and / or noise at the target spatial location due to the excessive distance. To address this issue, the microphone array can be positioned within the target area to minimize interference signals from speakers.

[0126] In some embodiments, the target region can be the region of minimum sound pressure level of the loudspeaker. The region of minimum sound pressure level can be the region where the loudspeaker radiates relatively little sound. In some embodiments, the loudspeaker can form at least one set of acoustic dipoles. For example, a set of sound signals with approximately opposite phases and approximately the same amplitude output from the front and back of the loudspeaker diaphragm can be regarded as two point sources. These two point sources can form an acoustic dipole or similar acoustic dipole, and the sound radiated outward has obvious directivity. Ideally, the sound radiated by the loudspeaker is larger in the straight line direction connecting the two point sources, and the sound radiated in other directions is significantly reduced, with the sound radiated by the loudspeaker being minimal in the region of the perpendicular bisector (or near the perpendicular bisector) of the line connecting the two point sources.

[0127] In some embodiments, the loudspeaker (e.g., loudspeaker 130) in the acoustic device (e.g., acoustic device 100) can be a bone conduction loudspeaker. When the loudspeaker is a bone conduction loudspeaker and the interference signal is the leakage signal of the bone conduction loudspeaker, the target area can be the region of minimum sound pressure level of the leakage signal of the bone conduction loudspeaker. The region of minimum sound pressure level of the leakage signal can refer to the region where the leakage signal radiated by the bone conduction loudspeaker is the smallest. The microphone array is positioned in the region of minimum sound pressure level of the leakage signal of the bone conduction loudspeaker, which can reduce the interference signal picked up by the microphone array from the bone conduction loudspeaker, and can also effectively solve the problem that the sound field of the target spatial location cannot be accurately estimated because the microphone array is too far away from the target spatial location.

[0128] Figure 11 This is a schematic diagram of the three-dimensional sound field leakage signal distribution of a bone conduction loudspeaker at 1000Hz, according to some embodiments of this application. Figure 12 This is a schematic diagram of the two-dimensional sound field leakage signal distribution of a bone conduction loudspeaker at 1000Hz, according to some embodiments of this application. Figure 11-12 As shown, the acoustic device 1100 may include a contact surface 1110. The contact surface 1110 may be configured to contact the user's body (e.g., face, ear) when the user wears the acoustic device 1100. A bone conduction speaker may be disposed within the acoustic device 1100. Figure 11 As shown, the colors on the acoustic device 1100 represent the sound leakage signal of the bone conduction speaker, and different color depths represent different levels of sound leakage. The lighter the color, the greater the sound leakage signal; the darker the color, the smaller the sound leakage signal. For example... Figure 11 As shown, the area 1120, where the dashed line is located, is darker in color compared to other areas, indicating a smaller sound leakage signal. Therefore, the area 1120 can be considered the region with the lowest sound pressure level of the sound leakage signal from the bone conduction speaker. As an example only, a microphone array could be positioned in the area 1120 (e.g., position 1) to receive a smaller sound leakage signal from the bone conduction speaker.

[0129] In some embodiments, the sound pressure level (SPL) of the minimum region of the bone conduction speaker can be 5-30 dB lower than the maximum output SPL of the bone conduction speaker. In some embodiments, the SPL of the minimum region of the bone conduction speaker can be 7-28 dB lower than the maximum output SPL of the bone conduction speaker. In some embodiments, the SPL of the minimum region of the bone conduction speaker can be 9-26 dB lower than the maximum output SPL of the bone conduction speaker. In some embodiments, the SPL of the minimum region of the bone conduction speaker can be 11-24 dB lower than the maximum output SPL of the bone conduction speaker. In some embodiments, the SPL of the minimum region of the bone conduction speaker can be 13-22 dB lower than the maximum output SPL of the bone conduction speaker. In some embodiments, the SPL of the minimum region of the bone conduction speaker can be 15-20 dB lower than the maximum output SPL of the bone conduction speaker. In some embodiments, the SPL of the minimum region of the bone conduction speaker can be 17-18 dB lower than the maximum output SPL of the bone conduction speaker. In some embodiments, the SPL of the minimum region of the bone conduction speaker can be 15 dB lower than the maximum output SPL of the bone conduction speaker.

[0130] Figure 12 The two-dimensional sound field distribution shown is Figure 11 A two-dimensional cross-sectional view of the three-dimensional sound field leakage signal distribution. (e.g.) Figure 12 As shown, the colors on the cross-section represent the sound leakage signal of the bone conduction speaker, and different color depths indicate different levels of sound leakage. Lighter colors indicate a larger sound leakage signal, while darker colors indicate a smaller sound leakage signal. Figure 12 As shown, regions 1210 and 1220, where the dashed lines are located, are darker in color compared to other regions, indicating lower sound leakage. Therefore, regions 1210 and 1220 can be considered the regions with the lowest sound pressure level of sound leakage from the bone conduction speaker. As an example only, a microphone array could be positioned in regions 1210 and 1220 (e.g., positions A and B) to receive lower sound leakage from the bone conduction speaker.

[0131] In some embodiments, the bone conduction loudspeaker emits a relatively large vibration signal during vibration. Therefore, not only does the sound leakage signal from the bone conduction loudspeaker interfere with the microphone array, but the vibration signal from the bone conduction loudspeaker also interferes with the microphone array. Here, the vibration signal of the bone conduction loudspeaker can refer to the vibration of other components of the acoustic device (e.g., housing, microphone array) caused by the vibration of the bone conduction loudspeaker's vibrating component. In this case, the interference signal from the bone conduction loudspeaker can include both the sound leakage signal and the vibration signal. To avoid the microphone array picking up the interference signal from the bone conduction loudspeaker, the target area where the microphone array is located can be the area with the minimum total energy of the sound leakage signal and vibration signal transmitted to the microphone array from the bone conduction loudspeaker. The sound leakage signal and vibration signal of the bone conduction loudspeaker are relatively independent signals, and the area with the minimum sound pressure level of the sound leakage signal of the bone conduction loudspeaker cannot represent the area with the minimum total energy of the sound leakage signal and vibration signal of the bone conduction loudspeaker. Therefore, determining the target area requires analysis of the total signal of the vibration signal and the sound leakage signal of the bone conduction loudspeaker.

[0132] Figure 13 This is a schematic diagram of the frequency response of the total signal of the vibration signal and the leakage signal of a bone conduction loudspeaker according to some embodiments of this application. Figure 13 The total signal of the bone conduction loudspeaker's vibration signal and leakage signal is shown in the figure. Figure 11 The frequency response curves at positions 1, 2, 3, and 4 on the acoustic device 1100. (Example) Figure 13 As shown, the horizontal axis represents frequency, and the vertical axis represents the sound pressure level of the total signal of the bone conduction loudspeaker's vibration signal and leakage signal. According to... Figure 11 According to the relevant description, when only considering the leakage signal of the bone conduction speaker, position 1, located in the area of ​​minimum sound pressure level of speaker 130, can be used as the target area for setting up the microphone array (e.g., microphone array 110, microphone array 820, microphone array 1020). However, when considering both the vibration signal and leakage signal of the bone conduction speaker, the target area for setting up the microphone array (i.e., the area where the total sound pressure level of the vibration signal and leakage signal of the bone conduction speaker is minimum) is not necessarily position 1. (Refer to...) Figure 13 Compared to other locations, the total sound pressure of the vibration signal and leakage signal of the bone conduction speaker at location 2 is relatively small. Therefore, location 2 can be used as the target area for setting up a microphone array.

[0133] In some embodiments, the location of the target area may be related to the orientation of the diaphragm of the microphone in the microphone array. The orientation of the microphone diaphragm can affect the magnitude of the bone conduction speaker vibration signal received by the microphone. For example, when the microphone diaphragm is perpendicular to the vibrating component of the bone conduction speaker, the microphone can collect a smaller bone conduction speaker vibration signal. Conversely, when the microphone diaphragm is parallel to the vibrating component of the bone conduction speaker, the microphone can collect a larger bone conduction speaker vibration signal. In some embodiments, the orientation of the microphone diaphragm can be set to reduce the bone conduction speaker vibration signal received by the microphone. For example, when the microphone diaphragm is perpendicular to the vibrating component of the bone conduction speaker, the bone conduction speaker vibration signal can be ignored during the determination of the target location of the microphone array, and only the leakage signal of the bone conduction speaker is considered, i.e., according to... Figure 11 and Figure 12 The description determines the target location for setting up the microphone array. For example, when the microphone diaphragm is parallel to the vibrating component of the bone conduction speaker, the determination of the target location for setting up the microphone array can simultaneously consider the vibration signal and leakage signal of the bone conduction speaker, i.e., based on... Figure 13 The description determines the target location for setting up the microphone array.

[0134] In some embodiments, adjusting the orientation of the microphone diaphragm can adjust the phase of the vibration signal from the bone conduction speaker received by the microphone, making the phase of the vibration signal from the bone conduction speaker received by the microphone approximately opposite to that of the leakage signal from the bone conduction speaker received by the microphone, and their magnitudes approximately equal. This allows the vibration signal from the bone conduction speaker received by the microphone to at least partially cancel out the leakage signal, thereby reducing the interference signal emitted by the bone conduction speaker received by the microphone array. In some embodiments, the vibration signal from the bone conduction speaker received by the microphone can reduce the leakage signal from the bone conduction speaker received by the microphone by 5-6 dB.

[0135] In some embodiments, the loudspeaker (e.g., loudspeaker 130) in the acoustic device (e.g., acoustic device 100) can be an air-conducting loudspeaker. When the loudspeaker is an air-conducting loudspeaker and the interference signal is the sound signal emitted by the air-conducting loudspeaker (i.e., the radiated sound field), the target area can be the region of minimum sound pressure level of the radiated sound field of the air-conducting loudspeaker. Positioning a microphone array within the region of minimum sound pressure level of the radiated sound field of the air-conducting loudspeaker can reduce the interference signal picked up by the microphone array from the air-conducting loudspeaker and effectively solve the problem of the microphone array being too far from the target spatial location, resulting in an inability to accurately estimate the sound field at the target spatial location.

[0136] Figure 14A -B is a schematic diagram of the sound field distribution of an air-conducting loudspeaker according to some embodiments of this application. For example... Figure 14AAs shown in -B, an air-conducting loudspeaker can be disposed within the open acoustic device 1400 and drawn from two sound-conducting holes (e.g., from the two sound-conducting holes of the open acoustic device 1400) Figure 14A -B (1401 and 1402) radiate sound outwards, and the emitted sound can form a dipole (with Figure 14A -B is indicated by the "+" and "-" symbols shown in the diagram.

[0137] like Figure 14A As shown, the open acoustic device 1400 is configured such that the line connecting the dipoles is approximately perpendicular to the user's face area. In this configuration, the sound radiated by the dipoles can form three relatively strong sound field regions 1421, 1422, and 1423. Between sound field regions 1421 and 1423, and between sound field regions 1422 and 1423, can form the region with the lowest sound pressure level (also referred to as the region with lower sound pressure) of the radiated sound field of the air-conducted loudspeaker. Figure 14A The area around the dashed line in FIG14 is considered the minimum sound pressure level region. This minimum sound pressure level region can refer to the area where the sound intensity output by the open acoustic device 1400 is relatively low. In some embodiments, the microphone 1430 in the microphone array can be located in this minimum sound pressure level region. For example, the microphone 1430 in the microphone array can be located at the position where the dashed line in FIG14 intersects with the housing of the open acoustic device 1400. This allows the microphone 1430 to receive as little sound signal as possible from the air-conducting speaker while collecting external ambient noise, thereby reducing the interference of the sound signal emitted by the air-conducting speaker on the active noise cancellation function of the open acoustic device 1400.

[0138] like Figure 14B As shown, the open-back acoustic device 1400 is configured such that the line connecting the dipoles is approximately parallel to the user's facial region. In this configuration, the sound radiated by the dipoles can form two relatively strong sound field regions (1424 and 1425). Between sound field regions 1424 and 1425, a region with minimal sound pressure level of the radiated sound field of the air-conducting loudspeaker can be formed, for example... Figure 14B The area shown is defined by the dashed line in Figure 14 and its surrounding area. In some embodiments, the microphone 1440 in the microphone array may be positioned in the area of ​​minimum sound pressure level. For example, the microphone 1440 in the microphone array may be positioned at the location where the dashed line in Figure 14 intersects with the housing of the open acoustic device 1400. This allows the microphone 1440 to receive as little sound signal as possible from the air-conducting speaker while collecting external ambient noise, thereby reducing the interference of the sound signal from the air-conducting speaker on the active noise cancellation function of the open acoustic device 1400.

[0139] Figure 15 This is an exemplary flowchart illustrating the output of a target signal based on a transfer function, according to some embodiments of this application. Figure 15 As shown, process 1500 may include:

[0140] In step 1510, the noise-reduced signal is processed based on the transfer function. In some embodiments, this step may be performed by processor 120 (e.g., amplitude-phase compensation unit 230). Further details regarding the noise-reduced signal can be found elsewhere in this application, for example... Figure 3 And its corresponding description. Additionally, according to Figure 3 As described, a loudspeaker (e.g., loudspeaker 130) can output a target signal based on a noise-reduced signal generated by processor 120.

[0141] In some embodiments, the target signal output by the speaker can be transmitted to a specific location in the user's ear (also known as the noise cancellation location) via a first acoustic path, and ambient noise can be transmitted to the same location via a second acoustic path. At this specific location, the target signal and ambient noise cancel each other out, so that the user cannot perceive the ambient noise or can only perceive a very weak ambient noise. In some embodiments, when the speaker is an air-conduction speaker, the specific location where the target signal and ambient noise cancel each other out can be the user's ear canal or its vicinity, for example, the target spatial location. The first acoustic path can be the path through which the target signal is transmitted from the air-conduction speaker to the target spatial location via the air, and the second acoustic path can be the path through which the ambient noise is transmitted from the noise source to the target spatial location. In some embodiments, when the speaker is a bone-conduction speaker, the specific location where the target signal and ambient noise cancel each other out can be the user's basilar membrane. The first acoustic path can be the path through which the target signal is transmitted from the bone-conduction speaker, through the user's bones or tissues, to the user's basilar membrane, and the second acoustic path can be the path through which the ambient noise is transmitted from the noise source, through the user's ear canal and tympanic membrane, to the user's basilar membrane.

[0142] In some embodiments, the loudspeaker (e.g., loudspeaker 130) may be positioned near the user's ear canal without obstructing it, thus maintaining a certain distance between the loudspeaker and the noise cancellation location (e.g., the target spatial location, the basilar membrane). Therefore, when the target signal output by the loudspeaker reaches the noise cancellation location, the phase and amplitude information of the target signal may change. As a result, the target signal output by the loudspeaker may fail to reduce ambient noise, or even amplify it, thereby preventing the active noise cancellation function of the acoustic device (e.g., the open-back acoustic output device 100) from being implemented.

[0143] Based on the above, the processor 120 can obtain the transfer function of the target signal from the speaker to the noise cancellation location. The transfer function may include a first transfer function and a second transfer function. The first transfer function may represent the changes in parameters of the target signal with respect to the acoustic path (i.e., the first acoustic path) from the speaker to the noise cancellation location (e.g., changes in amplitude and phase). In some embodiments, when the speaker is a bone conduction speaker, the target signal emitted by the bone conduction speaker is a bone conduction signal, and the location where the target signal emitted by the bone conduction speaker and the ambient noise are canceled is the user's basilar membrane. In this case, the first transfer function may represent the changes in parameters (e.g., phase and amplitude) of the target signal from the bone conduction speaker to the user's basilar membrane. In some embodiments, when the speaker is a bone conduction speaker, the first transfer function can be obtained experimentally. For example, the bone conduction speaker outputs the target signal, while simultaneously playing an air conduction sound signal with the same frequency as the target signal near the user's ear canal, and the cancellation effect of the target signal and the air conduction sound signal is observed. When the target signal and the air conduction sound signal cancel each other out, the first transfer function of the bone conduction speaker can be obtained based on the air conduction sound signal and the target signal output by the bone conduction speaker. In some embodiments, when the loudspeaker is an air-conducting loudspeaker, the signal emitted by the air-conducting loudspeaker to the target is an air-conducting sound signal, and the first transfer function can be obtained through acoustic diffusion field simulation and calculation. For example, the sound field of the target signal emitted by the air-conducting loudspeaker can be simulated using an acoustic diffusion field, and the first transfer function of the air-conducting loudspeaker can be calculated based on this sound field. The second transfer function can represent the changes in parameters of the ambient noise (e.g., changes in amplitude, changes in phase) from the target spatial location to the location where the target signal and ambient noise cancel each other out. As an example only, when the loudspeaker is a bone-conducting loudspeaker, the second transfer function can represent the changes in parameters of the ambient noise from the target spatial location to the user's basilar membrane. In some embodiments, the second transfer function can be obtained through acoustic diffusion field simulation and calculation. For example, the sound field of the ambient noise can be simulated using an acoustic diffusion field, and the second transfer function can be calculated based on this sound field.

[0144] In some embodiments, during the transmission of the target signal, not only will there be phase changes, but there may also be signal energy loss. Therefore, the transfer function may include a phase transfer function and an amplitude transfer function. In some embodiments, both the phase transfer function and the amplitude transfer function can be obtained using the methods described above.

[0145] Furthermore, the processor 120 can process the denoised signal based on the obtained transfer function. In some embodiments, the processor 120 can adjust the amplitude and phase of the denoised signal based on the obtained transfer function. In some embodiments, the processor 120 can adjust the phase of the denoised signal based on the obtained phase transfer function and adjust the amplitude of the denoised signal based on the amplitude transfer function.

[0146] In step 1520, a target signal is output based on the processed noise-reduced signal. In some embodiments, this step may be performed by the speaker 130.

[0147] In some embodiments, the speaker 130 can output a target signal based on the noise-reduced signal processed in step 1510, such that when the target signal output by the speaker 130 based on the processed noise-reduced signal reaches the location where the ambient noise is canceled out, the amplitude of the phase sum of the target signal and the ambient noise satisfies a specific condition. In some embodiments, the phase difference between the phase of the target signal and the phase of the ambient noise can be less than or equal to a certain phase threshold. This phase threshold can be in the range of 90-180 degrees. This phase threshold can be adjusted within this range according to the user's needs. For example, when the user does not want to be disturbed by the surrounding ambient sound, the phase threshold can be a larger value, such as 180 degrees, that is, the phase of the target signal is opposite to the phase of the ambient noise. As another example, when the user wants to remain sensitive to the surrounding environment, the phase threshold can be a smaller value, such as 90 degrees. It should be noted that the more ambient sound the user wants to receive, the closer the phase threshold can be to 90 degrees, and the less ambient sound the user wants to receive, the closer the phase threshold can be to 180 degrees. In some embodiments, when the phase of the target signal is constant with the phase of the ambient noise (e.g., out of phase), the amplitude difference between the amplitude of the ambient noise and the amplitude of the target signal can be less than or equal to a certain amplitude threshold. For example, when a user does not want to be disturbed by ambient noise, this amplitude threshold can be a small value, such as 0 dB, meaning the amplitude of the target signal is equal to the amplitude of the ambient noise. Alternatively, when a user wants to remain sensitive to their surroundings, this amplitude threshold can be a large value, such as approximately equal to the amplitude of the ambient noise. It should be noted that the more ambient noise a user wants to receive, the closer the amplitude threshold can be to the amplitude of the ambient noise; conversely, the less ambient noise a user wants to receive, the closer the amplitude threshold can be to 0 dB. This achieves the goal of reducing ambient noise and the active noise cancellation function of the acoustic device (e.g., acoustic output device 100), improving the user's auditory experience.

[0148] It should be noted that the above description of process 1500 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can make various modifications and changes to process 1500 under the guidance of this specification. For example, process 1500 may also include a step of obtaining the transfer function. As another example, steps 1510 and 1520 may be combined into one step. These modifications and changes are still within the scope of this application.

[0149] Figure 16 This is an exemplary flowchart illustrating noise estimation of a target spatial location according to some embodiments of this application. Figure 16 As shown, process 1600 may include:

[0150] In step 1610, components associated with the signal picked up by the bone conduction microphone are removed from the picked-up ambient noise in order to update the ambient noise.

[0151] In some embodiments, this step may be performed by processor 120. In some embodiments, when the microphone array (e.g., microphone array 110) picks up ambient noise, the user's own voice is also picked up by the microphone array; that is, the user's own voice is also considered part of the ambient noise. In this case, the target signal output by the speaker (e.g., speaker 130) will cancel out the user's own voice. In some embodiments, in certain scenarios, the user's own voice needs to be preserved, such as when the user is making a voice call or sending a voice message. In some embodiments, the acoustic device (e.g., acoustic device 100) may include a bone conduction microphone. When the user wears the acoustic device to make a voice call or record voice information, the bone conduction microphone can pick up the user's voice signal by picking up the vibration signal generated by the facial bones or muscles when the user speaks, and transmit it to processor 120. Processor 120 obtains parameter information of the voice signal picked up from the bone conduction microphone and removes the voice signal component associated with the voice signal picked up by the bone conduction microphone from the ambient noise picked up by the microphone array (e.g., microphone array 110). Processor 120 updates the ambient noise according to the parameter information of the remaining ambient noise. The updated ambient noise no longer includes the user's own voice signal, meaning that the user can hear the user's own voice signal when making a voice call.

[0152] In step 1620, the noise level of the target spatial location is estimated based on the updated ambient noise. In some embodiments, this step may be performed by processor 120. Step 1620 may be performed in a similar manner to step 320, and the relevant description will not be repeated here.

[0153] It should be noted that the above description of process 1600 is merely for illustration and explanation, and does not limit the scope of this application. Those skilled in the art can make various modifications and changes to process 1600 under the guidance of this application. For example, the components associated with the signal picked up by the bone conduction microphone can be preprocessed, and the signal picked up by the bone conduction microphone can be transmitted to the terminal device as an audio signal. These modifications and changes are still within the scope of this application.

[0154] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore remain within the spirit and scope of the exemplary embodiments of this application.

[0155] Furthermore, this application uses specific terms to describe its embodiments. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this application do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application can be appropriately combined.

[0156] Furthermore, those skilled in the art will understand that aspects of this application can be described and illustrated through several patentable types or situations, including any new and useful combination of processes, machines, products, or substances, or any new and useful improvements thereof. Accordingly, aspects of this application can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. All of the above hardware or software may be referred to as a “data block,” “module,” “engine,” “unit,” “component,” or “system.” Furthermore, aspects of this application may manifest as a computer product located on one or more computer-readable media, the product including computer-readable program code.

[0157] Computer storage media may contain a propagated data signal containing computer program code, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and suitable combinations thereof. Computer storage media can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer storage medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.

[0158] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this application are not intended to limit the order of the processes and methods of this application. Although the foregoing disclosure has discussed some currently considered useful embodiments of the invention through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the substance and scope of the embodiments of this application. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely through software solutions, such as installing the described system on existing servers or mobile devices.

[0159] Similarly, it should be noted that, in order to simplify the description of the present application and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of the embodiments of the present application sometimes combines multiple features into a single embodiment, drawing, or description thereof. However, this disclosure method does not imply that the subject matter of the application requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of the single embodiments disclosed above.

[0160] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of scope in some embodiments of this application are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0161] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this application, the entire contents of that patent are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this application, as well as documents that limit the broadest scope of the claims in this application (currently or subsequently appended to this application). It should be noted that if there are any inconsistencies or conflicts between the descriptions, definitions, and / or terminology used in the supplementary materials of this application and the content of this application, the descriptions, definitions, and / or terminology used in this application shall prevail.

[0162] Finally, it should be understood that the embodiments described in this application are merely illustrative of the principles of the embodiments of this application. Other modifications may also fall within the scope of this application. Therefore, alternative configurations of the embodiments of this application are considered as examples and not limitations, and are regarded as consistent with the teachings of this application. Accordingly, the embodiments of this application are not limited to the embodiments explicitly described and illustrated in this application.

Claims

1. An acoustic device, characterized in that, The acoustic device includes: The microphone array is configured to pick up ambient noise; The processor is configured as follows: The sound field at a target spatial location is estimated using the microphone array, where the target spatial location is closer to the user's ear canal than any microphone in the microphone array. A noise-reduced signal is generated based on the sound field estimation of the picked-up environmental noise and the target spatial location; and At least one speaker is configured to output a target signal based on the noise reduction signal, the target signal being used to reduce the ambient noise, wherein the microphone array is disposed in the target area to minimize interference signals from the at least one speaker to the microphone array; The generation of the denoised signal based on the sound field estimation of the picked-up environmental noise and the target spatial location includes: The noise level at the target's spatial location is estimated based on the acquired environmental noise; and The noise-reduced signal is generated based on the noise at the target spatial location and the sound field estimation at the target spatial location.

2. The acoustic device according to claim 1, characterized in that, The process of generating the denoised signal based on the noise at the target spatial location and the sound field estimation at the target spatial location includes: The difference between the physical quantity related to the sound field of the target spatial location and the parameter information of the noise of the target spatial location is determined, and the parameter information of the noise of the target spatial location is adjusted based on the difference; or the change in the parameter information of the environmental noise of the target spatial location is estimated based on the time information of the environmental noise picked up by the microphone array, the current time information, and the physical quantity related to the sound field of the target spatial location, and the parameter information of the noise of the target spatial location is adjusted based on the change. The noise reduction signal is generated based on the noise parameter information of the adjusted target spatial location, wherein the physical quantity includes sound.

3. The acoustic device according to claim 2, characterized in that, The acoustic device further includes one or more sensors for acquiring motion information of the acoustic device, and The processor is further configured to: The noise and sound field estimation of the target spatial location are updated based on the motion information. as well as The noise-reduced signal is generated based on the noise at the updated target spatial location and the sound field estimation at the updated target spatial location.

4. The acoustic device according to claim 2, characterized in that, The noise used to estimate the target spatial location based on the picked-up environmental noise includes: Identify one or more spatial noise sources related to the picked-up ambient noise; and Based on the spatial noise source, the noise at the target spatial location is estimated.

5. The acoustic device according to claim 1, characterized in that, The estimation of the sound field at the target spatial location using the microphone array includes: A virtual microphone is constructed based on the microphone array, the virtual microphone including a mathematical model or machine learning model, used to represent the audio data collected by the microphone if the target spatial location includes a microphone; and The sound field at the target spatial location is estimated based on the virtual microphone.

6. The acoustic device according to claim 5, characterized in that, The generation of the denoised signal based on the sound field estimation of the picked-up environmental noise and the target spatial location includes: The noise at the target spatial location is estimated based on the virtual microphone; and The noise-reduced signal is generated based on the noise at the target spatial location and the sound field estimation at the target spatial location.

7. The acoustic device according to claim 1, characterized in that, The at least one speaker is a bone conduction speaker. The interference signal includes the sound leakage signal and vibration signal of the bone conduction speaker, and The target area is the region where the total energy of the leakage signal and the vibration signal transmitted to the bone conduction speaker of the microphone array is minimized.

8. The acoustic device according to claim 1, characterized in that, The at least one loudspeaker is an air-conducting loudspeaker, and The target area is the region with the lowest sound pressure level in the radiated sound field of the air-conducting loudspeaker.

9. The acoustic device according to claim 1, characterized in that, The processor is further configured to process the noise-reduced signal based on a transfer function, the transfer function including a first transfer function and a second transfer function, the first transfer function representing the change of parameters of the target signal from the at least one loudspeaker to the location where the target signal and the ambient noise cancel each other out, and the second transfer function representing the change of parameters of the ambient noise from the target spatial location to the location where the target signal and the ambient noise cancel each other out. as well as The at least one speaker is further configured to output the target signal based on the processed noise-reduced signal.

10. A noise reduction method, characterized in that, The noise reduction method includes: Ambient noise is picked up by a microphone array; Processor The sound field of a target spatial location is estimated using the microphone array, where the target spatial location is closer to the user's ear canal than any microphone in the microphone array; A noise-reduced signal is generated based on the sound field estimation of the picked-up environmental noise and the target spatial location; and At least one loudspeaker outputs a target signal based on the noise reduction signal, the target signal being used to reduce the ambient noise, wherein the microphone array is positioned in the target area to minimize interference signals from the at least one loudspeaker to the microphone array; The generation of the denoised signal based on the sound field estimation of the picked-up environmental noise and the target spatial location includes: The noise level at the target's spatial location is estimated based on the acquired environmental noise; and The noise-reduced signal is generated based on the noise at the target spatial location and the sound field estimation at the target spatial location.

Citation Information

Patent Citations

  • Earphone device, headphone device, and method

    CN111095944A

  • Active noise reduction method and device, electronic equipment and chip

    CN111935589A

  • Acoustic device

    CN116918350A