Acoustic input and output devices

By designing the angle between the microphone vibration direction and the mechanical vibration direction of the speaker assembly in the acoustic input and output device, the problem of the impact of the speaker assembly on the microphone echo is solved, and the quality of the voice signal is improved.

CN115250392BActive Publication Date: 2025-05-23SHENZHEN SHOKZ CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110462049.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-27
Publication Date
2025-05-23
Estimated Expiration
2041-04-27

AI Technical Summary

Technical Problem

Mechanical vibrations of the speaker assembly are transmitted to the microphone, causing an increase in the intensity of the echo signal and reducing the quality of the microphone's voice signal collection.

Method used

By designing the vibration direction of the microphone and the mechanical vibration direction generated by the speaker assembly, the intensity of the echo signal is reduced and the quality of the voice signal is improved.

Benefits of technology

It effectively reduces the intensity of the echo signal received by the microphone, improves the quality of the voice signal, and makes the user experience better.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115250392B_ABST
    Figure CN115250392B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses an acoustic input-output device, comprising: a speaker assembly, used to transmit sound waves by generating a first mechanical vibration; and a microphone, used to receive a second mechanical vibration generated when a voice signal source provides a voice signal, the microphone generating a first signal and a second signal respectively under the action of the first mechanical vibration and the second mechanical vibration; a first angle formed by a vibration direction of the microphone and a direction of the first mechanical vibration is within a set angle range so that within a certain frequency range, the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of acoustics, and in particular to an acoustic input and output device. Background Art

[0002] The speaker assembly transmits sound by generating mechanical vibrations. The microphone receives the user's voice signal by picking up the vibrations of the user's skin and other parts when the user speaks. When the speaker assembly and the microphone work at the same time, the mechanical vibration of the speaker assembly will be transmitted to the microphone, causing the microphone to receive the vibration signal of the speaker assembly and generate an echo, which reduces the quality of the sound signal generated by the microphone and affects the user experience.

[0003] The present application provides an acoustic input and output device, which can reduce the influence of a speaker assembly on a microphone, reduce the strength of an echo signal generated by the microphone, and improve the quality of a voice signal collected by the microphone. Summary of the invention

[0004] The purpose of this application is to provide an acoustic input and output device, the purpose of which is to reduce the influence of the speaker assembly on the vibration of the bone conduction microphone, reduce the intensity of the echo signal generated by the bone conduction microphone, and improve the quality of the sound signal picked up by the bone conduction microphone.

[0005] In order to achieve the purpose of the above invention, the technical solution provided by this application is as follows:

[0006] An acoustic input-output device comprises: a speaker assembly for transmitting sound waves by generating a first mechanical vibration; and a microphone for receiving a second mechanical vibration generated when a voice signal source provides a voice signal, the microphone generating a first signal and a second signal respectively under the action of the first mechanical vibration and the second mechanical vibration; a first angle formed by a vibration direction of the microphone and a direction of the first mechanical vibration is within a set angle range so that within a certain frequency range, the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second signal.

[0007] In some embodiments, the first angle is in the range of 20 degrees to 90 degrees.

[0008] In some embodiments, the first angle is in the range of 75 degrees to 90 degrees.

[0009] In some embodiments, the first angle comprises 90 degrees.

[0010] In some embodiments, the second angle formed by the vibration direction of the microphone and the direction of the second mechanical vibration is within a set angle range so that the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second signal.

[0011] In some embodiments, the second angle is in the range of 0 degrees to 85 degrees.

[0012] In some embodiments, the second angle is in the range of 0 degrees to 15 degrees.

[0013] In some embodiments, the second angle includes 0 degrees.

[0014] In some embodiments, a vibration reduction structure is further included, the vibration reduction structure includes a vibration reduction material with an elastic modulus less than a first threshold, and the microphone is connected to the speaker assembly through the vibration reduction structure.

[0015] In some embodiments, the thickness of the vibration reduction structure is 0.5 mm to 5 mm. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The present application will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents a similar structure, wherein:

[0017] Figure 1 is a structural module diagram of an acoustic input and output device according to some embodiments of the present application;

[0018] Figure 2A and Figure 2B is a schematic diagram of the structure of an acoustic input and output device according to some embodiments of the present application;

[0019] Figure 3 is a schematic cross-sectional view of a partial structure of an acoustic input / output device according to some embodiments of the present application;

[0020] Figure 4 is a simplified schematic diagram of vibration transmission of an acoustic input-output device according to some embodiments of the present application;

[0021] Figure 5 is a schematic diagram of another mechanical vibration transmission model of an acoustic input-output device according to some embodiments of the present application;

[0022] Figure 6 is another structural schematic diagram of vibration transmission of an acoustic input-output device according to some embodiments of the present application;

[0023] Figure 7 is a schematic diagram of calculating and generating an electrical signal using a two-axis microphone according to some embodiments of the present application;

[0024] Figure 8 is a strength curve diagram of the second signal and the first signal according to some embodiments of the present application;

[0025] Fig. 9 is another intensity curve diagram of the second signal and the first signal according to some embodiments of the present application;

[0026] Fig.10 is a cross-sectional schematic diagram of the connection between a bone conduction microphone and a vibration reduction structure according to some embodiments of the present application;

[0027] Fig.11 is a cross-sectional schematic diagram of an acoustic input-output device with a vibration reduction structure according to some embodiments of the present application;

[0028] Fig.12 is a schematic cross-sectional view of an acoustic input / output device according to some embodiments of the present application;

[0029] Fig.13 is a schematic cross-sectional view of an acoustic input / output device according to some embodiments of the present application;

[0030] Fig.14 is a cross-sectional schematic diagram of an acoustic input-output device having two air conduction speaker assemblies according to some embodiments of the present application;

[0031] Fig.15 is another cross-sectional schematic diagram of an acoustic input-output device having two air conduction speaker assemblies according to some embodiments of the present application;

[0032] Fig.16 is a schematic diagram of the structure of a headset according to some embodiments of the present application;

[0033] Fig.17 is a schematic structural diagram of a single-ear headset according to some embodiments of the present application;

[0034] Fig.18 is a cross-sectional schematic diagram of a binaural headset according to some embodiments of the present application;

[0035] Fig.19 It is a schematic diagram of the structure of a pair of glasses according to some embodiments of the present application. DETAILED DESCRIPTION

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without paying creative work. It should be understood that these exemplary embodiments are given only to enable technicians in related fields to better understand and implement the present invention, and do not limit the scope of the present invention in any way. Unless it is obvious from the language environment or otherwise explained, the same reference numerals in the figures represent the same structure or operation.

[0037] As shown in the present application and claims, unless the context clearly indicates an exception, the words "a", "an", "a kind" and / or "the" do not specifically refer to the singular and may also include the plural, unless the context clearly indicates an exception. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements. The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment". The relevant definitions of other terms will be given in the following description. Below, without loss of generality, when describing the bone conduction related technology in the present invention, the description of "bone conduction microphone", "bone conduction microphone assembly", "bone conduction speaker", "bone conduction speaker assembly" or "bone conduction earphone" will be used. When describing the air conduction related technology in the present invention, the description of "air conduction microphone", "air conduction microphone assembly", "air conduction speaker", "air conduction speaker assembly" or "air conduction earphone" will be used. This description is only one form of bone conduction application. For ordinary technicians in this field, "device" or "headphones" can also be replaced by other similar words, such as "player", "hearing aid", etc. In fact, the various implementation methods in the present invention can be easily applied to other non-speaker devices. For example, for professionals in this field, after understanding the basic principle of the device, it is possible to make various modifications and changes in form and details to the specific methods and steps of implementing the device without deviating from this principle. In particular, the function of picking up and processing ambient sound is added to the device to enable the device to realize the function of a hearing aid. For example, a microphone such as a bone conduction microphone can pick up the sound of the user / wearer's surrounding environment, and transmit the processed sound (or the generated electrical signal) to the speaker component under a certain algorithm. That is, the bone conduction microphone can be modified to add the function of picking up ambient sound, and after certain signal processing, the sound is transmitted to the user / wearer through the speaker component, thereby realizing the function of a hearing aid. As an example, the algorithms mentioned here may include one or more combinations of noise elimination, automatic gain control, acoustic feedback suppression, wide dynamic range compression, active environment recognition, active noise reduction, directional processing, tinnitus processing, multi-channel wide dynamic range compression, active howling suppression, volume control, etc.

[0038] Figure 1 is a structural module diagram of an acoustic input and output device according to some embodiments of the present application. Figure 1 As shown, the acoustic input-output device 100 may include a speaker assembly 110 , a microphone assembly 120 , and a fixing assembly 130 .

[0039] The speaker assembly 110 can be used to convert a signal containing sound information into an acoustic signal (also referred to as a voice signal). For example, the speaker assembly 110 can generate mechanical vibrations to transmit sound waves (i.e., acoustic signals) in response to receiving a signal containing sound information. For the convenience of description, the mechanical vibrations generated by the speaker assembly 110 can be referred to as first mechanical vibrations. In some embodiments, the speaker assembly can include a vibration element and / or a vibration transmission element connected to the vibration element (e.g., at least a portion of the housing of the acoustic input / output device 100, a vibration transmission sheet). When the speaker assembly 110 generates the first mechanical vibration, it is accompanied by energy conversion, and the speaker assembly 110 can realize the conversion of a signal containing sound information into mechanical vibration. The conversion process may include the coexistence and conversion of multiple different types of energy. For example, an electrical signal (i.e., a signal containing sound information) can be directly converted into a first mechanical vibration through a transducer device in the vibration element of the speaker assembly 110, and the first mechanical vibration is conducted through the vibration transmission element of the speaker assembly 110 to transmit sound waves. For another example, the sound information can be contained in an optical signal, and a specific transducer device can realize the process of converting an optical signal into a vibration signal. Other energy types that can coexist and convert during the operation of the transducer include thermal energy, magnetic field energy, etc. The energy conversion methods of the transducer may include moving coil, electrostatic, piezoelectric, moving iron, pneumatic, electromagnetic, etc.

[0040] The speaker assembly 110 may include an air conduction speaker assembly and / or a bone conduction speaker assembly. In some embodiments, the speaker assembly 110 may include a vibration element and a housing. In some embodiments, when the speaker assembly 110 is a bone conduction speaker assembly, the housing of the speaker assembly 110 may be used to contact a part of the user's body (e.g., face) and transmit the first mechanical vibration generated by the vibration element to the auditory nerve via the bone, so that the user can hear the sound, and as at least part of the housing of the acoustic input-output device 100, the vibration element and the microphone assembly 120 are housed. In some embodiments, when the speaker assembly 110 is an air conduction speaker assembly, the vibration element may change the air density by pushing the air to vibrate, so that the user can hear the sound, and the housing may be at least part of the housing of the acoustic input-output device 100 to house the vibration element and the microphone assembly 120. In some embodiments, the speaker assembly 110 and the microphone assembly 120 may be located in different housings.

[0041] The vibration element can convert the sound signal into a mechanical vibration signal and thereby generate a first mechanical vibration. In some embodiments, the vibration element (i.e., the transducer) may include a magnetic circuit component. The magnetic circuit component may provide a magnetic field. The magnetic field may be used to convert a signal containing sound information into a mechanical vibration signal. In some embodiments, the sound information may include a video or audio file having a specific data format or data or a file that can be converted into sound through a specific path. The signal containing sound information may come from the storage component of the acoustic input-output device 100 itself, or from an information generation, storage or transmission system outside the acoustic input-output device 100. The signal containing sound information may include a combination of one or more electrical signals, optical signals, magnetic signals, mechanical signals, etc. The signal containing sound information may come from one signal source or multiple signal sources. Multiple signal sources may be related or unrelated. In some embodiments, the acoustic input-output device 100 may obtain a signal containing sound information in a variety of different ways, and the acquisition of the signal may be wired or wireless, and may be real-time or delayed. For example, the acoustic input-output device 100 may receive an electrical signal containing sound information in a wired or wireless manner, or may directly obtain data from a storage medium to generate a sound signal. For another example, the acoustic input-output device 100 may include a component with a sound collection function (e.g., an air conduction microphone component), which picks up the sound in the environment, converts the mechanical vibration of the sound into an electrical signal, and obtains the electrical signal that meets specific requirements after being processed by an amplifier. In some embodiments, the wired connection may include a metal cable, an optical cable, or a hybrid cable of metal and optical, such as a coaxial cable, a communication cable, a flexible cable, a spiral cable, a non-metallic sheathed cable, a metal sheathed cable, a multi-core cable, a twisted pair cable, a ribbon cable, a shielded cable, a telecommunication cable, a two-strand cable, a parallel two-core conductor, a twisted pair, or a combination of one or more thereof. The examples described above are for convenience of explanation only, and the medium of the wired connection may also be other types, such as a transmission carrier of other electrical signals or optical signals.

[0042] Wireless connection may include radio communication, free space optical communication, acoustic communication, and electromagnetic induction, etc. Radio communication may include IEEE802.11 series standards, IEEE802.15 series standards (such as Bluetooth technology and cellular technology, etc.), first generation mobile communication technology, second generation mobile communication technology (such as FDMA, TDMA, SDMA, CDMA, and SSMA, etc.), general packet radio service technology, third generation mobile communication technology (such as CDMA2000, WCDMA, TD-SCDMA, and WiMAX, etc.), fourth generation mobile communication technology (such as TD-LTE and FDD-LTE, etc.), satellite communication (such as GPS technology, etc.), near field communication (NFC) and other technologies running in ISM frequency bands (such as 2.4GHz, etc.); free space optical communication may include visible light, infrared signals, etc.; acoustic communication may include sound waves, ultrasonic signals, etc.; electromagnetic induction may include near field communication technology, etc. The examples described above are only for the convenience of explanation, and the medium of wireless connection may also be other types, such as Z-wave technology, other paid civil radio frequency bands and military radio frequency bands, etc. For example, as some application scenarios of the present technology, the acoustic input-output device 100 can obtain signals containing sound information from other acoustic input-output devices through Bluetooth technology.

[0043] The microphone assembly 120 can be used to pick up an acoustic signal (also referred to as a voice signal) and convert the acoustic signal into a signal containing sound information (e.g., an electrical signal). For example, the microphone assembly 120 picks up the mechanical vibration generated when the voice signal source provides a voice signal and converts it into an electrical signal. For the convenience of description, the mechanical vibration generated when the user provides a voice signal can be referred to as a second mechanical vibration. In some embodiments, the microphone assembly 120 may include one or more microphones. In some embodiments, the microphone can be divided into a bone conduction microphone and / or an air conduction microphone based on the working principle of the microphone. For the convenience of description, in one or more embodiments of the present application, a bone conduction microphone will be used as an example for explanation. It should be noted that the bone conduction microphone in one or more embodiments of the present application can also be replaced by an air conduction microphone.

[0044] The bone conduction microphone can be used to collect any mechanical vibrations (e.g., the first mechanical vibration and the second mechanical vibration) that are transmitted by the user's bones, skin, and other tissues and can be sensed by the bone conduction microphone. The received mechanical vibrations will cause the internal components of the bone conduction microphone 120 (e.g., the microphone diaphragm) to generate corresponding mechanical vibrations (e.g., the third mechanical vibration and the fourth mechanical vibration), and convert them into electrical signals containing voice information (e.g., the first signal and the second signal). The first signal can be understood as the echo signal generated by the bone conduction microphone; the second signal can be understood as the voice signal generated by the bone conduction microphone. The air conduction microphone can collect mechanical vibrations (i.e., sound waves) conducted by air and convert the mechanical vibrations into signals containing sound information (e.g., electrical signals). For example, if the speaker assembly 110 includes an air conduction speaker, the air conduction microphone can receive the echo signal transmitted by the air conduction speaker (transmitted by air conduction). For another example, if the speaker assembly 110 includes a bone conduction speaker, the air conduction microphone can simultaneously receive the mechanical vibration transmitted by the bone conduction speaker and the echo signal transmitted by the bone conduction speaker through the air conduction pathway. In some embodiments, the microphone assembly 120 may include a microphone diaphragm and other electronic components. After the mechanical vibration of the voice signal source is transmitted to the microphone diaphragm, the microphone diaphragm will generate corresponding mechanical vibrations. The electronic components can convert the mechanical vibration signal into a signal containing voice information (e.g., an electrical signal). In some embodiments, the microphone assembly 120 may include, but is not limited to, a ribbon microphone, a micro-electromechanical system (MEMS) microphone, a dynamic microphone, a piezoelectric microphone, a condenser microphone, a carbon microphone, an analog microphone, a digital microphone, etc., or any combination thereof. For another example, the bone conduction microphone may include an omnidirectional microphone, a unidirectional microphone, a bidirectional microphone, a cardioid microphone, etc., or any combination thereof.

[0045] In some embodiments, when the speaker assembly 110 and the microphone assembly 120 work simultaneously, the microphone assembly 120 can sense the first mechanical vibration generated by the speaker assembly 110 and the second mechanical vibration generated by the voice signal source. In response to the first mechanical vibration, the microphone assembly 120 can generate a third mechanical vibration and convert the third mechanical vibration into a first signal. In response to the second mechanical vibration, the microphone assembly 120 can generate a fourth mechanical vibration and convert the fourth mechanical vibration into a second signal. In some embodiments, the speaker assembly 110 can be referred to as an echo signal source. In some embodiments, when the speaker assembly 110 and the microphone assembly 120 work simultaneously, within a certain frequency range, the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second. The frequency range may include 200Hz to 10kHz, or 200Hz to 5000Hz, or 200Hz to 2000Hz, or 200Hz to 1000Hz, etc.

[0046] The fixing assembly 130 can support the speaker assembly 110 and the microphone assembly 120. In some embodiments, the fixing assembly 130 can include an arc-shaped elastic component that can form a force that rebounds toward the middle of the arc so that it can be in stable contact with the human skull. In some embodiments, the fixing assembly 130 can include one or more connectors. One or more connectors can connect the speaker assembly 110 and / or the microphone assembly 120. In some embodiments, the fixing assembly 130 can be worn in both ears. For example, the two ends of the fixing assembly 130 can be fixedly connected to two groups of speaker assemblies 110 respectively. When the user wears the acoustic input and output device 100, the fixing assembly 130 can fix the two groups of speaker assemblies 110 near the left and right ears of the user respectively. In some embodiments, the fixing assembly 130 can also be worn in one ear. For example, the fixing assembly 130 can be fixedly connected to only one group of speaker assemblies 110. When the user wears the acoustic input and output device 100, the fixing assembly 130 can fix the speaker assembly 110 near one ear of the user. In some embodiments, the fixing component 130 can be any combination of one or more of glasses (e.g., sunglasses, augmented reality glasses, virtual reality glasses), a helmet, a headband, etc., without limitation herein.

[0047] The above description of the structure of the acoustic input and output device is only a specific example and should not be regarded as the only feasible implementation scheme. Obviously, for professionals in this field, after understanding the basic principle of the acoustic input and output device 100, it is possible to make various modifications and changes in form and details to the specific methods and steps of implementing the acoustic input and output device 100 without deviating from this principle, but these modifications and changes are still within the scope of the above description. For example, the acoustic input and output device 100 may include one or more processors, and the processor may execute one or more sound signal processing algorithms. The sound signal processing algorithm can modify or enhance the sound signal. For example, the sound signal is subjected to noise reduction, acoustic feedback suppression, wide dynamic range compression, automatic gain control, active environment recognition, active anti-noise, directional processing, tinnitus processing, multi-channel wide dynamic range compression, active howling suppression, volume control, or other similar or any combination of the above processing, and these modifications and changes are still within the scope of protection of the claims of the present invention. For another example, the acoustic input and output device 100 may include one or more sensors, such as a temperature sensor, a humidity sensor, a speed sensor, a displacement sensor, etc. The sensor can collect user information or environmental information.

[0048] Figure 2A and Figure 2B Schematic diagram of the structure of the acoustic input and output device shown in some embodiments of the present application. Figure 2A and Figure 2BAs shown, in some embodiments, the acoustic input and output device 200 can be an ear clip type earphone, and the ear clip type earphone can include an earphone core 210, a fixing assembly 230, a control circuit 240 and a battery 250. The earphone core 210 can include a speaker assembly (not shown in the figure) and a microphone assembly (not shown in the figure). The fixing assembly can include an ear hook 231, an earphone housing 232, a circuit housing 233 and a back hang 234. The earphone housing 232 and the circuit housing 233 can be respectively arranged at the two ends of the ear hook 231, and the back hang 234 can be further arranged at the end of the circuit housing 233 that is farther from the ear hook 231. The earphone housing 232 can be used to accommodate different earphone cores. The circuit housing 233 can be used to accommodate a control circuit 260 and a battery 270. The two ends of the back hang 234 can be respectively connected to the corresponding circuit housing 233. The ear hook 231 may refer to a structure for hanging the ear clip-type earphone on the user's ear when the user wears the acoustic input and output device 200, and fixing the earphone housing 232 and the earphone core 210 at a predetermined position relative to the user's ear.

[0049] In some embodiments, the ear hook 231 may include an elastic metal wire. The elastic metal wire may be configured to keep the ear hook 231 in a shape that matches the user's ear and has a certain elasticity, so that when the user wears the ear clip earphone, it can undergo a certain elastic deformation according to the user's ear shape and head shape to adapt to users with different ear shapes and head shapes. In some embodiments, the elastic metal wire may be made of a memory alloy with good deformation recovery ability. Even if the ear hook 231 is deformed by an external force, it may return to its original shape when the external force is removed, thereby extending the service life of the ear clip earphone. In some embodiments, the elastic metal wire may also be made of a non-memory alloy. A conductor may be provided in the elastic metal wire to establish an electrical connection between the earphone core 210 and other components (such as a control circuit 260, a battery 270, etc.) to facilitate providing power and data transmission for the earphone core 210. In some embodiments, the ear hook 231 may also include a protective sleeve 236 and a shell protector 237 formed integrally with the protective sleeve 236.

[0050] In some embodiments, the earphone housing 232 may be configured to accommodate the earphone core 210. The earphone core 210 may include one or more speaker assemblies and / or one or more microphone assemblies. The one or more speaker assemblies may include a bone conduction speaker assembly, an air conduction speaker assembly, etc. The one or more microphone assemblies may include a bone conduction microphone assembly, an air conduction microphone assembly, etc. The structure and arrangement of the speaker assembly and the microphone assembly may refer to the description elsewhere in this application, for example, Figure 3-Figure 15 The number of the earphone core 210 and the earphone housing 232 can be two, which can correspond to the left ear and the right ear of the user respectively.

[0051] In some embodiments, the ear hook 231 and the earphone housing 232 can be molded separately and further assembled together, rather than directly molding the two together.

[0052] In some embodiments, the earphone housing 232 may be provided with a contact surface 2321. The contact surface 2321 may be in contact with the user's skin. When using an ear clip-type earphone, the sound waves generated by one or more bone conduction speakers of the earphone core 210 may be transferred to the outside of the earphone housing 232 (for example, to the eardrum of the user) through the contact surface 221. In some embodiments, the material and thickness of the contact surface 2321 may affect the propagation of bone conduction sound waves to the user, thereby affecting the sound quality. For example, if the material of the contact surface 2321 is more elastic, the transmission of bone conduction sound waves in the low frequency range may be better than the transmission of bone conduction sound waves in the high frequency range. On the contrary, if the material of the contact surface 2321 is less elastic, the transmission of bone conduction sound waves in the high frequency range may be better than the transmission of bone conduction sound waves in the low frequency range. It should be noted that the earphone housing 232 in this implementation and the housing in other embodiments of the present application are used to refer to the parts of the acoustic input and output device 200 that contact the user.

[0053] Figure 3 Schematic diagram of a cross-section of a portion of the structure of an acoustic input / output device shown in some embodiments of the present application. Figure 3 As shown, in some embodiments, the acoustic input-output device 300 may include a speaker assembly 310, which may be used to transmit sound waves by generating a first mechanical vibration; and a bone conduction microphone 320, which may be used to receive a second mechanical vibration generated when a voice signal source provides a voice signal. In some embodiments, the acoustic input-output device 300 may also include a fixing assembly 330, such as Figure 3As shown, the fixing component 330 is fixedly connected to the speaker component 310. When the user wears the acoustic input and output device 300, the speaker component 310 and the bone conduction microphone 320 are kept in contact with the user's face 340. In some embodiments, when the bone conduction microphone 320 and the speaker component 310 work simultaneously, the bone conduction microphone 320 can receive the first mechanical vibration and the second mechanical vibration, respectively generate the third mechanical vibration and the fourth mechanical vibration under the action of the first mechanical vibration and the second mechanical vibration, and convert the third mechanical vibration and the fourth mechanical vibration into the first signal and the second signal respectively. In some embodiments, within a certain frequency range, the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second signal. As described herein, the third mechanical vibration can also be referred to as the first mechanical vibration received by the bone conduction microphone 320, that is, the echo signal received by the bone conduction microphone 320; the fourth mechanical vibration can also be referred to as the second mechanical vibration received by the bone conduction microphone 320, that is, the voice signal received by the bone conduction microphone 320. In some embodiments, the frequency range can include 200Hz to 10kHz. In some embodiments, the frequency range may include 200 Hz to 9000 Hz. In some embodiments, the frequency range may include 200 Hz to 8000 Hz. In some embodiments, the frequency range may include 200 Hz to 6000 Hz. In some embodiments, the frequency range may include 200 Hz to 5000 Hz.

[0054] The speaker assembly 310 can transmit sound waves by generating a first mechanical vibration so that the user can hear the sound. The speaker assembly 310 transmits sound waves in two ways: air conduction and bone conduction. Among them, the air conduction speaker assembly corresponds to the transmission of sound waves through air conduction. The air conduction speaker assembly propagates sound waves through the air in the form of waves, and the sound waves are transmitted to the auditory nerve via the user's eardrum-auditory ossicles-cochlea, so that the user can hear the sound. The bone conduction speaker assembly corresponds to the transmission of sound waves through bone conduction. The bone conduction speaker assembly transmits mechanical vibrations to the skin and bones of the user's face 340 and transmits them to the auditory nerve through the bones by contacting the user's face 340 (for example, the shell 350 of the bone conduction speaker assembly contacts the user's face 340), so that the user can hear the sound. Whether it is a bone conduction speaker assembly or an air conduction speaker assembly, the bone conduction microphone 320 will be directly or indirectly connected to the speaker assembly 310. Specifically, when the speaker assembly 310 is a bone conduction speaker assembly, the housing 350 is one of the vibration transmission elements of the bone conduction speaker assembly. The vibration element in the bone conduction speaker assembly needs to be directly or indirectly connected to the housing 350 so as to transmit the vibration to the user's skin and bones. The bone conduction microphone 320 needs to be directly or indirectly connected to the housing 350 to collect the vibration generated when the user speaks. When the bone conduction speaker transmits sound waves, it will cause mechanical vibration of the housing 350, and the housing 350 will transmit the mechanical vibration to the bone conduction microphone 320. After receiving the mechanical vibration, the bone conduction microphone 320 will generate a corresponding third mechanical vibration and generate a first signal containing sound information based on the third mechanical vibration. When the speaker assembly 310 is an air conduction speaker assembly, the housing 350 is used to accommodate the air conduction speaker assembly and the bone conduction microphone 320, which is equivalent to the outer shell of the acoustic input and output device 300. The vibration element in the air conduction speaker assembly can be directly or indirectly connected to the housing 350 to fix the air conduction speaker assembly. In summary, the bone conduction microphone 320 needs to be directly or indirectly connected to the housing 350 to collect the vibration generated when the user speaks. When the air conduction speaker transmits sound waves, it will cause mechanical vibration of the shell 350, and the shell 350 will transmit the mechanical vibration to the bone conduction microphone 320. After receiving the mechanical vibration, the bone conduction microphone 320 will generate a corresponding third mechanical vibration and generate a first signal containing sound information based on the third mechanical vibration.

[0055] Therefore, at least a part of the first mechanical vibration generated by the speaker assembly 310 is transmitted to the bone conduction microphone 320, causing the bone conduction microphone 320 to generate a third mechanical vibration. In addition to the first mechanical vibration transmitted by the speaker assembly 310, the bone conduction microphone 320 can receive the second mechanical vibration (e.g., vibration of the skin and bones) generated when the user speaks by contacting the skin of the user's face 340, causing the bone conduction microphone 320 to generate a fourth mechanical vibration.

[0056] When the bone conduction microphone 320 and the speaker assembly 310 work at the same time, for example, the bone conduction microphone 320 receives a voice signal (for example, by picking up the vibration of the skin or other parts of the body when the person speaks, receiving the voice signal of the person speaking) and the speaker assembly 310 transmits a voice signal (for example, music) through vibration, the bone conduction microphone 320 will receive the first mechanical vibration and the second mechanical vibration at the same time. The microphone diaphragm (not shown in the figure) of the bone conduction microphone 320 will generate a third mechanical vibration and a fourth mechanical vibration corresponding to the first mechanical vibration and the second mechanical vibration, respectively, and will convert the third mechanical vibration and the fourth mechanical vibration into a first signal and a second signal, respectively. When the microphone diaphragm generates the third mechanical vibration in response to the first mechanical vibration picked up, the bone conduction microphone 320 will receive the voice information transmitted by the first mechanical vibration in addition to the voice information transmitted by the second mechanical vibration, thereby affecting the quality of the sound signal picked up by the microphone. For the convenience of description, the signal transmitted by the first mechanical vibration can be referred to as an echo signal (or a secondary voice signal), and the component that generates and transmits the first mechanical vibration (for example, the speaker assembly 310, the housing 350) can be referred to as an echo signal source (or a secondary voice signal source). The second mechanical vibration can be called a voice signal (or a main voice signal), and the component that generates and transmits the second mechanical vibration (for example, the user's vocal cords, nasal cavity, mouth, etc.) can be called a voice signal source (or a main voice signal source). Figure 3 The vibration directions of the voice signal source, the echo signal source and the bone conduction microphone are shown, wherein the direction indicated by arrow A is the direction of the first mechanical vibration, that is, the vibration direction of the echo signal source; the direction indicated by arrow B is the vibration direction of the bone conduction microphone, that is, the direction of the third mechanical vibration and the fourth mechanical vibration; the direction indicated by arrow C is the direction of the second mechanical vibration, that is, the vibration direction of the voice signal source.

[0057] Based on the above reasons, it is necessary to reduce the intensity of the echo signal (i.e., the intensity of the first signal) generated by the bone conduction microphone 320 by designing the acoustic input and output device 300. Furthermore, while reducing the intensity of the echo signal generated by the bone conduction microphone 320, the intensity of the voice signal (i.e., the intensity of the second signal) generated by the bone conduction microphone 320 can be increased, thereby achieving the purpose of reducing the intensity of the first signal and increasing the intensity of the second signal, so that the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second signal, thereby improving the quality of the sound signal generated by the bone conduction microphone.

[0058] Figure 4 Schematic diagram of vibration transmission of acoustic input and output devices shown in some embodiments of the present application. Figure 3 and Figure 4As shown, when the bone conduction microphone 320 and the speaker assembly 310 in the acoustic input-output device 300 work simultaneously, the mechanical vibration transmission model of the acoustic input-output device 300 can be equivalent to Figure 4 The model shown. Specifically, the strength of the mechanical vibration (i.e., the second mechanical vibration) of the voice signal source 360 ​​(e.g., the user's bones or vocal cords) is L1; the strength of the mechanical vibration (i.e., the first mechanical vibration) of the echo signal source 380 (e.g., the speaker assembly 310) is L2; ​​the bone conduction microphone 320 and the voice signal source 360 ​​may be connected by a first elastic connection 370, and the elastic coefficient of the first elastic connection 370 is k1; the bone conduction microphone 320 and the echo signal source 380 may be connected by a second elastic connection 390, and the elastic coefficient of the second elastic connection 390 is k2; the mass of the bone conduction microphone 320 is m. The first elastic connection 370 between the voice signal source 360 ​​and the bone conduction microphone 320 may include the contact parts between the bone conduction microphone 320 and the user's face 340 (e.g., the vibration transmission layer, the metal sheet, part of the housing 350, etc.), the user's skin, etc. The second elastic connection 390 between the bone conduction microphone 320 and the echo signal source 380 is a part of the acoustic input and output device 300. For example, the bone conduction microphone 320 and the echo signal source 380 may be physically connected to the housing 350 at the same time, and the second elastic connection 390 may include the housing 350. For another example, the bone conduction microphone 320 and the echo signal source 380 may be physically connected to the housing 350 via a connector, respectively, and the second elastic connection 390 may include the housing 350 and the connector. Figure 4 In the illustrated embodiment, it can be assumed that the vibration direction of the voice signal source 360 ​​is parallel to the vibration direction of the bone conduction microphone 320, and the vibration direction of the echo signal source 380 is parallel to the vibration direction of the bone conduction microphone 320, so that the bone conduction microphone can receive the vibration of the voice signal source 360 ​​and the vibration of the echo signal source 380 to the greatest extent. The vibration direction of the bone conduction microphone 320 can be understood as the direction of vibration of the microphone diaphragm.

[0059] according to Figure 4 , the intensity L of the mechanical vibration received by the bone conduction microphone 320 can be obtained as:

[0060]

[0061] Wherein, L1 is the intensity of the second mechanical vibration received by the bone conduction microphone 320 (i.e., the fourth mechanical vibration intensity), L2 is the intensity of the first mechanical vibration received (i.e., the third mechanical vibration intensity), and m is the mass of the bone conduction microphone 320. ω is the angular frequency of the signal, and the signal includes a voice signal and / or an echo signal. It can represent the influence of L1 (i.e. the second mechanical vibration) on L; The influence of L2 (ie the first mechanical vibration) on L can be represented.

[0062] From this, it can be seen that the larger the elastic coefficient k1 of the first elastic connection 370, the greater the influence of the vibration intensity L1 of the voice signal source 360 ​​on the intensity L of the mechanical vibration received by the bone conduction microphone 320; the smaller the elastic coefficient k2 of the second elastic connection 390, the smaller the influence of the vibration intensity L2 of the echo signal source 380 on the intensity L of the mechanical vibration received by the bone conduction microphone 320, and the smaller the echo signal received by the bone conduction microphone 320.

[0063] Based on formula (1), it can be known that in order to reduce the echo signal received by the bone conduction microphone 320, the acoustic input and output device can be designed from multiple aspects. For example, L1 and / or k1 can be increased as much as possible, and L2 and / or k2 can be reduced as much as possible to increase the influence of L1 on L and reduce the influence of L2 on L, thereby improving the quality of the sound signal generated by the bone conduction microphone.

[0064] Figure 5 Schematic diagram of another mechanical vibration transmission model of the acoustic input and output device shown in some embodiments of the present application. Figure 5 As shown, in some embodiments, the bone conduction microphone 520 may be a uniaxial bone conduction microphone, and the microphone diaphragm of the uniaxial bone conduction microphone can only vibrate in one direction, that is, the microphone diaphragm can only convert the mechanical vibration in this direction into an electrical signal (e.g., the first signal). Figure 5 For example, the vibration direction of the bone conduction microphone 520 is the up-down direction. When the direction of the mechanical vibration is parallel to the vibration direction of the bone conduction microphone 520 (i.e., both are in the up-down direction), the microphone diaphragm can convert the received mechanical vibration into an electrical signal (e.g., the first signal and the second signal) to the greatest extent. Here, converting the received mechanical vibration into an electrical signal to the greatest extent can be understood as that almost all mechanical vibrations except for the loss caused by resistance and the like (e.g., a part of the mechanical vibration will be lost when it is transmitted through the first elastic connection 570 and the second elastic connection 590) can be received by the microphone diaphragm and converted into an electrical signal. When the direction of the mechanical vibration is perpendicular to the vibration direction of the bone conduction microphone 520 (i.e., the left-right direction), only a small part of the received mechanical vibration can be converted into an electrical signal by the microphone diaphragm, so the intensity of the electrical signal is the smallest. That is to say, when the vibration direction of the bone conduction microphone 520 is perpendicular to the direction of the mechanical vibration, the intensity of the electrical signal generated by the bone conduction microphone 520 is the smallest, and the intensity of the generated sound signal is the smallest.

[0065] Based on the above principle, in some embodiments, the installation position of the bone conduction microphone 520 can be designed so that the vibration direction of the bone conduction microphone 520 is aligned with the echo signal source 580 (for example, Figure 3 The vibration direction of the speaker assembly 310 (shown in FIG. 1 ) (i.e., the first mechanical vibration direction) is within a certain angle range to reduce the strength of the first signal generated by the bone conduction microphone 520, that is, to reduce the strength of the echo signal generated by the bone conduction microphone 520. Further, in some embodiments, the vibration direction of the bone conduction microphone 520 is aligned with the voice signal source 560 (e.g., Figure 3 The vibration direction of the user's face 340) is within a certain angle range to increase the strength of the second signal generated by the bone conduction microphone 520, that is, to increase the strength of the voice signal generated by the bone conduction microphone 520.

[0066] Figure 6 is another structural schematic diagram of the vibration transmission of the acoustic input and output device shown in some embodiments of the present application. Figure 6 As shown, in some embodiments, the vibration direction of the bone conduction microphone 620 is consistent with the echo signal source 680 (eg, Figure 3 The angle formed by the vibration direction of the loudspeaker assembly 310 shown in the figure can be a first angle α. In some embodiments, the first angle α can be in the angle range of 20 degrees to 90 degrees. In some embodiments, the first angle α can be in the angle range of 45 degrees to 90 degrees. In some embodiments, the first angle α can be in the angle range of 60 degrees to 90 degrees. In some embodiments, the first angle α can be in the angle range of 75 degrees to 90 degrees. In some embodiments, the first angle α can be 90 degrees. In this embodiment, within the range of 20 degrees to 90 degrees, the larger the angle of the first angle α, the closer the vibration direction of the microphone diaphragm is to the vibration direction of the echo signal source 680, and the smaller the intensity of the first signal converted by the microphone diaphragm. When the first angle α is 90 degrees, the intensity of the first signal converted by the microphone diaphragm is the smallest, that is, the intensity of the echo signal generated by the bone conduction microphone 620 is the smallest.

[0067] In some embodiments, according to formula (1), it can be known that the greater the influence of the vibration intensity L1 of the voice signal source 660 on the intensity L of the mechanical vibration received by the bone conduction microphone 620, that is, the greater the vibration intensity L1 of the voice signal source 660 received by the bone conduction microphone 620, the smaller the influence of the vibration intensity L2 of the echo signal source 680 on the intensity L of the mechanical vibration received by the bone conduction microphone 620. In some embodiments, in order to increase the influence of the vibration intensity L1 of the voice signal source 660 on the sound signal L generated by the bone conduction microphone 620, the angle between the vibration direction of the bone conduction microphone 620 and the vibration direction of the voice signal source 660 can be designed to be within a certain range. Among them, the angle between the vibration direction of the bone conduction microphone 620 and the vibration direction of the voice signal source 660 can be a second angle β. In some embodiments, the second angle β can be within an angle range of 0 degrees to 85 degrees. In some embodiments, the second angle β can be within an angle range of 0 degrees to 75 degrees. In some embodiments, the second angle β can be within an angle range of 0 degrees to 60 degrees. In some embodiments, the second angle β can be within an angle range of 0 degrees to 45 degrees. In some embodiments, the second angle β may be in the angle range of 0 to 30 degrees. In some embodiments, the second angle β may be in the angle range of 0 to 15 degrees. In some embodiments, the second angle β may be in the angle range of 0 to 5 degrees. In some embodiments, the second angle β may be 0 degrees, that is, the vibration direction of the bone conduction microphone 620 is parallel to the vibration direction of the voice signal source 660. In this embodiment, within the range of 0 to 90 degrees, the smaller the second angle β is, the closer the vibration direction of the microphone diaphragm is to the vibration direction of the voice signal source 660, and the greater the intensity of the second signal converted by the microphone diaphragm. When the second angle β is 0 degrees, the intensity of the first signal converted by the microphone diaphragm is the largest, and at this time, the intensity of the second signal generated by the bone conduction microphone 620 is the largest, that is, the intensity of the generated voice signal is the largest. As described herein, the angle between two directions refers to the smallest positive angle formed by the intersection of the straight lines in which the two directions are located.

[0068] It should be noted that the scheme of controlling the first angle α within a set angle range can be combined with the scheme of controlling the second angle β within a set angle range. In some embodiments, the first angle α can be set to 90 degrees, and the second angle β can be set to 30 degrees. In some embodiments, the first angle α can be set to 90 degrees, and the second angle β can be set to 45 degrees. In some embodiments, the first angle α can be set to 90 degrees, and the second angle β can be set to 60 degrees. In some embodiments, the first angle α can be set to 45 degrees, and the second angle β can be set to 30 degrees. In some embodiments, the first angle α can be set to 90 degrees, and the second angle β can be set to 15 degrees. When the first angle α is set to 90 degrees and the second angle β is set to 0 degrees, Figure 6 and Figure 5 In this embodiment, the bone conduction microphone 620 can convert the vibration received from the voice signal source 660 into the second signal to the greatest extent, and the intensity of the generated first signal is minimized, thereby improving the quality of the sound signal generated by the bone conduction microphone 620.

[0069] Figure 8 It is a strength curve diagram of the second signal and the first signal shown in some embodiments of the present application. Figure 8 shows a bone conduction microphone based on Figure 4 The mechanical vibration (i.e., the first mechanical vibration) generated by the echo signal source 380 and the intensity curve 810 of the first signal and the intensity curve 820 of the second signal converted based on the mechanical vibration (i.e., the second mechanical vibration) generated by the voice signal source 360, wherein the horizontal axis is the frequency and the vertical axis is the sound intensity. In some embodiments, Figure 8 The first signal and the second signal strength curves shown are obtained when the first angle α is 0 degrees and the second angle β is also 0 degrees. Figure 3 , Figure 4 and Figure 8 It can be seen that within the frequency range of about 0 to 500 Hz, the strength of the first signal generated by the bone conduction microphone 320 is less than the strength of the second signal. When the frequency exceeds 500 Hz, for example, within the frequency range of 500 Hz to 10000 Hz, the strength of the first signal generated by the bone conduction microphone 320 is greater than the strength of the second signal, and the echo generated by the bone conduction microphone 320 is relatively large. Therefore, the strength of the echo signal generated by the bone conduction microphone 320 can be reduced by designing the installation position of the bone conduction microphone 320 and the speaker assembly 310.

[0070] For example, Fig. 9 is another intensity curve diagram of the first signal and the second signal shown in some embodiments of the present application. Fig. 9 As shown, in this embodiment, the bone conduction microphone 620 and the echo signal source 680 (for example, Figure 3 The position of the speaker assembly 310 shown in the figure is designed so that the first angle α is 90 degrees and the second angle β is 60 degrees. It can be seen from the first signal intensity curve 810 and the first signal intensity curve 910 and the second signal intensity curve 820 and the second signal intensity curve 920 that after the above design (i.e., adjusting the first angle α and the second angle β), the intensity of the first signal generated by the bone conduction microphone 620 is significantly reduced (e.g., Fig. 9At the same time, the above design has little or almost negligible effect on the weakening of the second signal strength generated by the bone conduction microphone 620, and the intensity of the first signal strength reduction generated by the bone conduction microphone 620 is significantly less than the intensity of the first signal strength reduction, so that the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second signal. In some embodiments, after adopting the above design, within the frequency range of 0 to 800 Hz, the intensity of the first signal generated by the bone conduction microphone 620 is smaller than that of Figure 8 In general, within a wider low-frequency range, the intensity of the first signal generated by the bone conduction microphone 620 is smaller, that is, the intensity of the echo signal generated by the bone conduction microphone 620 is smaller, so that the user can hear a clearer voice signal, effectively improve the sound quality, and effectively improve the user experience.

[0071] In some embodiments, after designing the positions of the bone conduction microphone 620 and the echo signal source 680 (for example, the speaker assembly 310), the amplitude of the decrease in the intensity of the second signal is significantly smaller than the amplitude of the decrease in the intensity of the first signal, so that the ratio of the intensity of the second signal to the intensity of the first signal can be greater than the threshold, thereby increasing the proportion of the voice signal in the sound signal generated by the bone conduction microphone 620, making the voice signal clearer and the user experience better. In some embodiments, the ratio of the intensity of the second signal to the intensity of the first signal can be greater than 1 / 4. In some embodiments, the ratio of the intensity of the second signal to the intensity of the first signal can be greater than 1 / 3. In some embodiments, the ratio of the intensity of the second signal to the intensity of the first signal can be greater than 1 / 2. In some embodiments, the ratio of the intensity of the second signal to the intensity of the first signal can be greater than 2 / 3.

[0072] It should be noted that the microphone assembly (for example, Figure 3 The scheme of reducing the strength of the echo signal can also be applied to the air conduction microphone.

[0073] In some embodiments, a single-axis bone conduction microphone is described only as an example. In addition, a bone conduction microphone (e.g., Figure 3 The bone conduction microphone 320 shown may also be other types of microphones. For example, the bone conduction microphone 320 may be a two-axis microphone, a three-axis microphone, a vibration sensor, an accelerometer, etc.

[0074] Continue to refer Figure 3 and Figure 4In some embodiments, the bone conduction microphone 320 may be a two-axis microphone, that is, the bone conduction microphone 320 may convert mechanical vibrations received in two directions into electrical signals. Figure 7 Schematic diagram of the two-axis microphone calculation and generation of electrical signals according to some embodiments of the present application. In some embodiments, the two directions may have a certain angle (i.e., a third angle). The angle range of the third angle is 0 degrees to 90 degrees. Figure 7 As shown, the two directions are represented as the X-axis direction and the Y-axis direction, and the X-axis is perpendicular to the Y-axis. The angle between the echo signal source 380 and the X-axis of the bone conduction microphone is α(e), the angle between the speech signal source 360 ​​and the X-axis of the bone conduction microphone is β(s), the echo signal (i.e., the first mechanical vibration) generated by the echo signal source 380 is e(t), and the speech signal (i.e., the second mechanical vibration) generated by the speech signal source 360 ​​is s(t), then the vibration components of the echo signal source 380 and the speech signal source 360 ​​on the X-axis of the bone conduction microphone are:

[0075] x(t)=e(t)cos(α(e))+s(t)cos(β(s)), (2)

[0076] The vibration components of the echo signal source 380 and the voice signal source 360 ​​on the Y-axis of the bone conduction microphone are:

[0077] y(t)=e(t)sin(α(e))+s(t)sin(β(s)), (3)

[0078] The echo signal of the bone conduction microphone 320 can be eliminated by weighting the vibration component x(t) of the echo signal source 380 and the voice signal source 360 ​​on the X-axis of the bone conduction microphone and the vibration component y(t) of the echo signal source 380 and the voice signal source 360 ​​on the Y-axis of the bone conduction microphone. Then, the total sound signal of the bone conduction microphone 320 is:

[0079] out(t)=x(t)sin(α(e))-y(t)cos(d(e))=s(t)sin(α(e)-β(s)), (4)

[0080] Among them, the weighting coefficient corresponding to the vibration component x(t) of the echo signal source 380 and the voice signal source 360 ​​on the X-axis of the bone conduction microphone is sin(α(e)), and the weighting coefficient corresponding to the vibration component y(t) of the echo signal source 380 and the voice signal source 360 ​​on the Y-axis of the bone conduction microphone is -cos(α(e)). In some embodiments, the angle α(e) between the echo signal source 380 and the X-axis of the bone conduction microphone can be obtained when the acoustic input and output device is assembled. In some embodiments, α(e) can be obtained by the following process, including determining whether the current signal of the bone conduction microphone 320 has a voice signal s(t); when the current signal does not have a voice signal s(t), the size of α(e) is obtained by the following formulas (5)-(7).

[0081] x(t)=e(t)cos(α(e)), (5)

[0082] y(t)=e(t)sin(α(e)), (6)

[0083] According to formulas (5) and (6), we can get:

[0084]

[0085] In some embodiments, x(t) and y(t) may be weighted and α(e) may be obtained according to formula (7). In some embodiments, after solving α(e) according to formula (9), α(e) may be smoothed in time to obtain a more stable α(e) estimate.

[0086] In some embodiments, the bone conduction microphone 320 may also be a three-axis microphone. For example, the microphone may have an X-axis, a Y-axis, and a Z-axis, and the sound signal generated by the three-axis microphone may be obtained by weighted calculation based on the components of the speech signal s(t) and the echo signal e(t) on the X-axis, Y-axis, and Z-axis of the bone conduction microphone. Since the principle of calculating and generating a sound signal by a three-axis microphone is similar to that of a two-axis microphone, it will not be described in detail here.

[0087] In some embodiments, the vibration direction of the echo signal source 380 may not be a single direction, for example, the vibration direction of the echo signal source 380 may be diffused along an arc trajectory. In this case, the vibration generated by the echo signal source 380 that is not perpendicular to the vibration direction of the bone conduction microphone 320 can be received by the bone conduction microphone 320 and converted into a first signal, that is, an echo signal is generated. Therefore, in some embodiments, the speaker assembly 310 and the bone conduction microphone 320 can be designed so that the position between the bone conduction microphone 320 and the speaker assembly 310 (for example, the housing 350) is relatively fixed to reduce the vibration transmitted by the echo signal source 380 received by the bone conduction microphone 320.

[0088] In some embodiments, in addition to designing the first angle α and the second angle β, the purpose of reducing echo may also be achieved by changing the elastic coefficient k1 of the first elastic connection 370 and the elastic coefficient k2 of the second elastic connection 390 .

[0089] In some embodiments, the intensity of the first mechanical vibration (ie, the third mechanical vibration) received by the bone conduction microphone 320 may be reduced by reducing the elastic strength k2 of the second elastic connection 390 between the bone conduction microphone 320 and the echo signal source 380 .

[0090] Fig.10 is a cross-sectional schematic diagram of the connection between the bone conduction microphone and the vibration reduction structure shown in some embodiments of the present application, Fig.11 Schematic diagram of a cross section of an acoustic input / output device with a vibration reduction structure shown in some embodiments of the present application. Fig.10 and Fig.11 As shown, the acoustic input-output device 1000 may include a bone conduction microphone 1020 and a speaker assembly 1010. The bone conduction microphone 1020 and the speaker assembly 1010 may be placed in the same housing. In some embodiments, the acoustic input-output device 1000 may further include a vibration reduction structure 1100, and the bone conduction microphone 1020 may be connected to the speaker assembly 1010 through the vibration reduction structure 1100. When the bone conduction microphone 1020 and the speaker assembly 1010 work simultaneously, the speaker assembly 1010 may transmit a voice signal (sound wave) through a first mechanical vibration, and the bone conduction microphone 1020 may receive or transmit a second mechanical vibration generated when a voice signal source provides a voice signal to pick up the voice signal. The first mechanical vibration of the speaker assembly 1010 may be transmitted to the bone conduction microphone 1020 through the vibration reduction structure 1100, and the bone conduction microphone 1020 may generate a third mechanical vibration and a fourth mechanical vibration under the action of the first mechanical vibration and the second mechanical vibration. The vibration reduction structure 1100 can reduce the intensity of the first mechanical vibration of the speaker assembly 1010 (echo signal source) received by the bone conduction microphone 1020 , thereby reducing the intensity of the first signal generated by the bone conduction microphone 1020 .

[0091] The vibration reduction structure 1100 may refer to a structure with a certain elasticity, and the strength of the mechanical vibration transmitted from the echo signal source 1080 is reduced by its elasticity. In some embodiments, the vibration reduction structure 1100 may be an elastic member to reduce the strength of the transmitted mechanical vibration. The elasticity of the vibration reduction structure 1100 may be determined by the material, thickness, structure, etc. of the vibration reduction structure.

[0092] In some embodiments, the vibration damping structure 1100 may be made of a vibration damping material having an elastic modulus less than a first threshold. In some embodiments, the first threshold may be 5000 MPa. In some embodiments, the first threshold may be 4000 MPa. In some embodiments, the first threshold may be 3000 MPa. In some embodiments, the elastic modulus of the vibration damping material may be in the range of 0.01 MPa to 1000 MPa. In some embodiments, the elastic modulus of the vibration damping material may be in the range of 0.015 MPa to 2500 MPa. In some embodiments, the elastic modulus of the vibration damping material may be in the range of 0.02 MPa to 2000 MPa. In some embodiments, the elastic modulus of the vibration damping material may be in the range of 0.025 MPa to 1500 MPa. In some embodiments, the elastic modulus of the vibration damping material may be in the range of 0.03 MPa to 1000 MPa. In some embodiments, the vibration damping material may include, but is not limited to, foam, plastic (for example, but not limited to high molecular weight polyethylene, blown nylon, engineering plastics, etc.), rubber, silicone, etc. In some embodiments, the vibration damping material may be foam.

[0093] In some embodiments, the vibration reduction structure 1100 may have a certain thickness. Fig.10 As shown, the thickness of the vibration reduction structure 1100 can be understood as the dimension in any one of the X-axis direction, the Y-axis direction or the Z-axis direction. In some embodiments, the thickness of the vibration reduction structure 1100 can be in the range of 0.5 mm to 5 mm. In some embodiments, the thickness of the vibration reduction structure 1100 can be in the range of 1 mm to 4.5 mm. In some embodiments, the thickness of the vibration reduction structure 1100 can be in the range of 1.5 mm to 4 mm. In some embodiments, the thickness of the vibration reduction structure 1100 can be in the range of 2 mm to 3.5 mm. In some embodiments, the thickness of the vibration reduction structure 1100 can be in the range of 2 mm to 3 mm.

[0094] In some embodiments, the elasticity of the vibration reduction structure 1100 can be provided by its structural design. For example, the vibration reduction structure 1100 can be an elastic structure, and even if the material used to make the vibration reduction structure 1100 has a high rigidity, the elasticity can be provided by its structure. In some embodiments, the vibration reduction structure 1100 can include but is not limited to a spring-like structure, a ring-shaped or ring-like structure, etc.

[0095] In some embodiments, the surface of the bone conduction microphone 1020 may include a first portion 1021 and a second portion 1022, wherein the first portion 1021 may be used to contact the user's face 1040 to conduct the second mechanical vibration provided by the voice signal source, and the second portion 1022 may be used to connect with other components of the acoustic input and output device 1000 (for example, connected with the speaker assembly 1010), and the second portion 1022 may be provided with a vibration reduction structure 1100, and then connected with the speaker assembly 1010 through the vibration reduction structure 1100. In this embodiment, the vibration reduction structure 1100 provided between the speaker assembly 1010 and the bone conduction microphone 1020 has a certain elasticity, which can reduce the first mechanical vibration transmitted by the speaker assembly 1010, reduce the intensity of the first mechanical vibration received by the bone conduction microphone 1020, and make the echo signal generated by the bone conduction microphone 1020 smaller. Furthermore, the reason why the vibration reduction structure 1100 is not provided on the first part 1021 is that the first part 1021 of the surface of the bone conduction microphone 1020 is in contact with the user's face 1040 to conduct the second mechanical vibration. For example, the first part 1021 may be a side close to the microphone diaphragm, and the second mechanical vibration represents the voice signal provided by the voice signal source, so it is necessary to ensure that the second mechanical vibration is not weakened as much as possible. Specifically, in combination with Fig.10 and Fig.11 As shown, the vibration reduction structure 1100 can surround the second portion 1022 of the surface of the bone conduction microphone 1020 and leave the first portion 1021 empty so that the first portion 1021 can directly contact the user's face 1040.

[0096] In some embodiments, the vibration reduction structure 1100 can be connected to the second portion 1022 of the surface of the bone conduction microphone by adhesive. In some embodiments, the vibration reduction structure 1100 can also be fixed to the bone conduction microphone 1020 by welding, clamping, riveting, threaded connection (for example, connected by screws, screws, screws, bolts, etc.), clamp connection, pin connection, wedge key connection, or integrated molding.

[0097] In some embodiments, the first portion 1021 of the surface of the bone conduction microphone 1020 may be provided with a vibration transmission layer 1023. Since the bone conduction microphone 1020 has a large rigidity, if the first portion 1021 directly contacts the user's face 1040, the user may feel uncomfortable, which will reduce the user experience. After the vibration transmission layer 1023 is provided on the first portion 1021, the touch is better when in contact with the user, which can effectively improve the user experience.

[0098] In some embodiments, the vibration transmission layer 1023 needs to maintain a certain elasticity, which can not only reduce the loss of the second mechanical vibration during the conduction process, but also ensure that the user has a good touch after wearing the acoustic input and output device 1000. In some embodiments, if the elastic modulus of the material of the vibration transmission layer 1023 is too small, it means that the elasticity of the material of the vibration transmission layer 1023 is small, which will weaken the intensity of the second mechanical vibration. Therefore, in some embodiments, the elastic modulus of the material used to make the vibration transmission layer 1023 can be greater than the second threshold. In some embodiments, the second threshold can be 0.01Mpa. In some embodiments, the second threshold can be 0.015Mpa. In some embodiments, the second threshold can be 0.02Mpa. In some embodiments, the second threshold can be 0.025Mpa. In some embodiments, the second threshold can be 0.03Mpa. In some embodiments, the elastic modulus of the vibration transmission layer 1023 can be in the range of 0.03MPa to 3000MPa. In some embodiments, the elastic modulus of the vibration transmission layer 1023 can be in the range of 5MPa to 2000MPa. In some embodiments, the elastic modulus of the vibration transmission layer 1023 may be in the range of 10 MPa to 1500 MPa. In some embodiments, the elastic modulus of the vibration transmission layer 1023 may be in the range of 10 MPa to 1000 MPa. In some embodiments, the material for making the vibration transmission layer 1023 may be silicone (the elastic modulus of silicone is 10 MPa), rubber or plastic (the elastic modulus of plastic is 1000 MPa).

[0099] In some embodiments, the loss of the second mechanical vibration during the conduction process can be reduced by reducing the thickness of the vibration transmission layer 1023. When the thickness of the vibration transmission layer 1023 is thin, even if the elastic modulus of the material used to make the vibration transmission layer 1023 is small, the intensity of the second mechanical vibration will not be greatly lost. In some embodiments, the thickness of the vibration transmission layer 1023 can be less than 30 mm. In some embodiments, the thickness of the vibration transmission layer 1023 can be less than 25 mm. In some embodiments, the thickness of the vibration transmission layer 1023 can be less than 20 mm. In some embodiments, the thickness of the vibration transmission layer 1023 can be less than 15 mm. In some embodiments, the thickness of the vibration transmission layer 1023 can be less than 10 mm. In some embodiments, the thickness of the vibration transmission layer 1023 can be less than 5 mm. In some embodiments, the vibration transmission layer 1023 can be made of rubber or silicone with a thickness of 5 mm, which can ensure a good touch while also ensuring the intensity of the second mechanical vibration received by the bone conduction microphone 1020.

[0100] It should be noted that the above-described embodiments of the acoustic input-output device 1000 are applicable to both bone conduction speaker assemblies and air conduction speaker assemblies. For example, when it is a bone conduction speaker assembly, the housing 1050 may be a part of the bone conduction speaker assembly, and the bone conduction microphone 1020 may be connected to the housing of the bone conduction speaker assembly through a vibration reduction structure 1100. When it is an air conduction speaker assembly, the air conduction speaker assembly and the bone conduction microphone 1020 may both be connected to the housing (for example, the diaphragm is connected to the housing, and the bone conduction microphone 1020 is connected to the housing), and a vibration reduction structure is also provided between the bone conduction microphone 1020 and the housing.

[0101] In some embodiments, the intensity of the second mechanical vibration (i.e., the fourth mechanical vibration) received by the bone conduction microphone can be increased by increasing the clamping force on the part of the acoustic input-output device 1000 in contact with the user. It can be understood that the closer the contact between the acoustic input-output device 1000 and the user's contact part (e.g., the user's face 1040), the less loss the second mechanical vibration will suffer during the transmission process. However, if the clamping force on the part of the acoustic input-output device 1000 in contact with the user is large, the user will feel pain and have a poor user experience. Therefore, it is necessary to control the clamping force within a certain range. In some embodiments, when the speaker assembly 1010 is an air conduction speaker assembly, that is, the acoustic input-output device 1000 transmits a sound signal to the user through the air conduction speaker assembly, and receives the user's voice signal through the bone conduction microphone 1020, in this case, the clamping force can be set within the range of 0.001N to 0.3N. In some embodiments, the clamping force can be set within the range of 0.0025N to 0.25N. In some embodiments, the clamping force can be set within the range of 0.005N to 0.15N. In some embodiments, the clamping force may be set within a range of 0.0075N to 0.1N. In some embodiments, the clamping force may be set within a range of 0.01N to 0.05N. In some embodiments, since the bone conduction speaker assembly transmits the mechanical vibration generated by the vibration element to the user's face via the housing so that the user can hear the sound, when the speaker assembly 1010 is a bone conduction speaker assembly, the clamping force is different. For example, when the speaker assembly 1010 of the acoustic input-output device 1000 includes a bone conduction speaker assembly, if the clamping force is too small, the intensity of the mechanical vibration transmitted to the user by the bone conduction speaker assembly will also be too small, that is, the volume of the sound transmitted to the user by the acoustic input-output device 1000 is too small. Therefore, in order to ensure the intensity of the mechanical vibration received by the user, in some embodiments, when the speaker assembly 1010 of the acoustic input-output device 1000 includes a bone conduction speaker assembly, the clamping force needs to be set within a certain range. In some embodiments, the clamping force may be set within a range of 0.01N to 2.5N. In some embodiments, the clamping force may be set within a range of 0.025N to 2N. In some embodiments, the clamping force may be set within a range of 0.05 N to 1.5 N. In some embodiments, the clamping force may be set within a range of 0.075 N to 1 N. In some embodiments, the clamping force may be set within a range of 0.1 N to 0.5 N.

[0102] In some embodiments, the speaker assembly 1010 and the bone conduction microphone 1020 may be directly connected, for example, the bone conduction microphone 1020 is directly connected to the housing 1050 (housing of the bone conduction speaker assembly) of the speaker assembly 1010 and is accommodated in the housing 1050. In some embodiments, the bone conduction microphone and the speaker assembly may be indirectly connected.

[0103] Fig.12 1 is a cross-sectional schematic diagram of an acoustic input / output device shown in some embodiments of the present application. In some embodiments, the acoustic input / output device 1200 includes a speaker assembly 1210 and a bone conduction microphone 1220. The speaker assembly 1210 is a bone conduction speaker assembly. The speaker assembly 1210 may include a housing 1250 and a vibration element 1211 connected to the housing 1250 for generating a first mechanical vibration in transmitting sound waves. The bone conduction microphone 1220 is connected to the housing 1250. Fig.12 As shown, the vibration element 1211 may include a vibration transmitting piece 1213, a magnetic circuit assembly 1215 and a coil 1217 (or a voice coil). The magnetic circuit assembly 1215 may be used to form a magnetic field, and the coil 1217 may generate mechanical vibrations in the magnetic field, thereby causing the vibration transmitting piece 1213 to vibrate. Specifically, when a signal current is passed through the coil 1217, the coil 1217 is in the magnetic field formed by the magnetic circuit assembly 1215, and generates mechanical vibrations under the action of the Ampere force. The vibration of the coil 1217 drives the vibration transmitting piece 1213 to generate mechanical vibrations. And the mechanical rotation of the vibration transmitting piece 1213 may be further transferred to the housing 1250, and then contact the user through the housing 1250 so that the user can hear the sound.

[0104] In some embodiments, the bone conduction microphone 1220 can be disposed at any position on the inner wall of the housing 1250, for example, Fig.12 The inner wall of the lower side of the housing 1250 is shown as being connected to the inner wall on the left side. For another example, the inner wall disposed on the lower side of the housing 1250 does not contact the inner wall on the left side or the right side. The acoustic input-output device 1200 can be combined with one or more of the above-mentioned embodiments, for example, in Fig.12 A vibration reduction structure is provided between the bone conduction microphone 1220 and the housing 1250 to reduce the intensity of the first mechanical vibration received by the bone conduction microphone 1220 .

[0105] Fig.131 is a cross-sectional schematic diagram of an acoustic input-output device shown in some embodiments of the present application. The acoustic input-output device 1300 includes a speaker assembly 1310 and a bone conduction microphone 1320. In some embodiments, the speaker assembly 1310 is an air conduction speaker assembly, and the speaker assembly 1310 may include a housing 1350 and a vibration element 1311. The vibration element 1311 may include a diaphragm 1313, a magnetic circuit assembly 1315, and a coil 1317. The magnetic circuit assembly 1315 may be used to form a magnetic field, and the coil 1317 may mechanically vibrate in the magnetic field to cause vibration of the diaphragm 1313. There is a first connection between the housing 1350 and the vibration element 1311. The first connection may include a first vibration reduction structure.

[0106] When the air conduction speaker assembly is working, the diaphragm 1313 will generate mechanical vibration, and because the diaphragm 1313 and the housing 1350 are directly connected (such as Fig.13 As shown in FIG. 1 , the vibration of the diaphragm 1313 causes mechanical vibration of the housing 1350. Fig.12 The difference between the bone conduction speaker assembly shown is that the air conduction speaker assembly does not need to rely on the vibration of the housing 1350 to transmit sound waves, but relies on a number of sound-transmitting holes (e.g., the first sound-transmitting hole 1351 and the second sound-transmitting hole 1352) opened on the housing to transmit sound waves to the user. Therefore, a first vibration reduction structure can be provided between the vibration element 1311 and the housing 1350 to reduce the mechanical vibration of the housing 1350, thereby reducing the intensity of the mechanical vibration transmitted by the housing 1350 and received by the bone conduction microphone 1320.

[0107] In some embodiments, the first vibration damping structure may be arranged in the same or similar manner as the vibration damping structure 1100 in the aforementioned embodiment. For example, the first vibration damping structure may be made of the same thickness, the same material, and the same structure as the vibration damping structure 1100. In some embodiments, the first vibration damping structure may be different from the vibration damping structure 1100. For example, the first vibration damping structure may be a strip-shaped member or a sheet-shaped member with a certain elasticity. The two ends of the strip-shaped member or the sheet-shaped member are respectively connected to the diaphragm 1313 and the shell 1350 to reduce the intensity of the mechanical vibration transmitted from the diaphragm 1313 to the shell 1350. The first vibration damping structure may also be an annular member. The middle part of the annular member is connected to the diaphragm, and the outer side of the annular member is connected to the shell 1350, which can also reduce the intensity of the mechanical vibration transmitted from the diaphragm 1313 to the shell 1350.

[0108] Continue to refer Fig.13 In some embodiments, a second connection may be included between the housing 1350 and the bone conduction microphone 1320. The second connection may include a second vibration reduction structure. The second vibration reduction structure may reduce the intensity of the mechanical vibration (i.e., the third mechanical vibration) transmitted to the bone conduction microphone 1320 via the housing 1350.

[0109] In some embodiments, the bone conduction microphone 1320 and the speaker assembly 1310 can be respectively arranged in different areas of the acoustic input-output device, and then a second vibration reduction structure is arranged between the bone conduction microphone 1320 and the housing 1350 of the speaker assembly 1310. In some embodiments, the bone conduction microphone 1320 can be separately arranged in other areas of the acoustic input-output device, and then connected to the housing 1350 through the second vibration reduction structure. Fig.17 Taking the embodiment shown as an example, the acoustic input and output device 1700 is a single-ear headset, and the bone conduction microphone 1720 and the speaker assembly 1710 are respectively arranged in two earmuffs 1731 on both sides of the fixing assembly 1730, and then connected through the fixing assembly 1730. Fig.17 In the embodiment shown, the second connection includes a fixing component 1730 and earmuffs 1731 disposed on both sides of the fixing component 1730. A second vibration reduction structure may be disposed on the fixing component 1730 and the earmuffs 1731. For example, a layer of vibration reduction material is disposed on the outer surface of the fixing component 1730 as the second vibration reduction structure. Fig.18 In the embodiment shown, the acoustic input-output device 1800 is a binaural headset, the earmuff 1831 is provided with a sponge cover 1833, the bone conduction microphone 1820 is provided in the sponge cover 1833, and is connected to the housing 1850 of the speaker assembly 1810 through the sponge cover 1833. In this embodiment, the sponge cover 1833 can be equivalent to a second vibration reduction structure, reducing the intensity of the first mechanical vibration transmitted to the bone conduction microphone 1820. For a detailed description of the second vibration reduction structure, please refer to other embodiments of the present application (such as Fig.17 , Fig.18 and Fig.19 ), which will not be described in detail here.

[0110] The above-mentioned embodiment of the second vibration reduction structure is applicable not only to air conduction speaker assemblies, but also to bone conduction speaker assemblies. Fig.17 , Fig.18 The speaker assembly in the illustrated embodiment may be replaced by Fig.12 The bone conduction speaker assembly shown. Fig.17 For example, the bone conduction speaker assembly and the bone conduction microphone 1720 are respectively arranged in two earmuffs 1731, and a layer of vibration-damping material can still be set on the fixing assembly 1730 as a second vibration-damping structure.

[0111] It should be noted that when the bone conduction microphone is Fig.13 As shown, the second vibration reduction structure is arranged inside the housing, and when the bone conduction microphone is directly connected to the housing, the second vibration reduction structure is the same as the vibration reduction structure in the aforementioned embodiment. For more description, please refer to Fig.10 and Fig.11 The relevant content will not be repeated here.

[0112] refer to Fig.13 As shown, in some embodiments, the mechanical vibration intensity of the housing 1350 can be reduced not only by adding a first vibration reduction structure between the vibration element 1311 and the housing 1350, but also by other means. In some embodiments, the influence of the vibration element 1311 on the housing 1350 when vibrating can be reduced by reducing the mass of the vibration element 1311, thereby reducing the mechanical vibration intensity of the housing 1350. The vibration element 1311 can include a diaphragm 1313, and the mechanical vibration of the housing 1350 is caused by the vibration of the diaphragm 1313. If the mass of the vibration element 1311 (for example, the diaphragm 1313) is small, the influence of the vibration element 1311 on the housing 1350 when vibrating will be reduced, and the intensity of the mechanical vibration generated by the housing 1350 will be small. In some embodiments, the mass of the diaphragm 1313 can be controlled within the range of 0.001g to 1g. In some embodiments, the mass of the diaphragm 1313 can be controlled within the range of 0.002g to 0.9g. In some embodiments, the mass of the diaphragm 1313 can be controlled within a range of 0.003g to 0.8g. In some embodiments, the mass of the diaphragm 1313 can be controlled within a range of 0.004g to 0.7g. In some embodiments, the mass of the diaphragm 1313 can be controlled within a range of 0.005g to 0.6g. In some embodiments, the mass of the diaphragm 1313 can be controlled within a range of 0.005g to 0.5g. In some embodiments, the mass of the diaphragm 1313 can be controlled within a range of 0.005g to 0.3g.

[0113] Similarly, if the mass of the housing 1350 is much greater than the mass of the diaphragm 1313, the mechanical vibration of the diaphragm 1313 has less effect on the housing 1350. Therefore, in some embodiments, the strength of the mechanical vibrator of the housing 1350 can be reduced by increasing the mass of the housing 1350. In some embodiments, the mass of the housing 1350 can be controlled within the range of 2g to 20g. In some embodiments, the mass of the housing 1350 can be controlled within the range of 3g to 15g. In some embodiments, the mass of the housing 1350 can be controlled within the range of 4g to 10g. In some embodiments, the ratio of the mass of the housing 1350 to the mass of the diaphragm 1313 can be controlled so that the mass of the housing 1350 is much greater than the mass of the diaphragm 1313, thereby reducing the effect of the mechanical vibration of the diaphragm 1313 on the housing 1350. In some embodiments, the ratio of the mass of the housing 1350 to the mass of the diaphragm 1313 can be controlled within the range of 10 to 100. In some embodiments, the ratio of the mass of the housing 1350 to the mass of the diaphragm 1313 can be controlled within the range of 15 to 80. In some embodiments, the ratio of the mass of the housing 1350 to the mass of the diaphragm 1313 can be controlled within a range of 20 to 60. In some embodiments, the ratio of the mass of the housing 1350 to the mass of the diaphragm 1313 can be controlled within a range of 25 to 50. In some embodiments, the ratio of the mass of the housing 1350 to the mass of the diaphragm 1313 can be controlled within a range of 30 to 50.

[0114] Fig.14 is a cross-sectional schematic diagram of an acoustic input-output device having two air conduction speaker assemblies shown in some embodiments of the present application, Fig.15 It is a cross-sectional schematic diagram of another acoustic input-output device having two air conduction speaker components shown in some embodiments of the present application. Fig.14 and Fig.15 In the embodiment shown, the speaker components are all air conduction speaker components. Fig.14 As shown, in some embodiments, the speaker assembly 1410 may include a first vibration element 1411 and a second vibration element 1412, the first vibration element 1411 includes a first diaphragm 1413, a first magnetic circuit assembly 1415 and a first coil 1417, and the second vibration element 1412 includes a second diaphragm 1414, a second magnetic circuit assembly 1416 and a second coil 1418 (or a voice coil). In some embodiments, the vibration directions of the first diaphragm 1413 and the second diaphragm 1414 are opposite. For example, Fig.14The vibration direction of the first diaphragm 1413 and the second diaphragm 1414 at a certain moment is shown, wherein the vibration direction of the first diaphragm 1413 is from top to bottom, and the vibration direction of the second diaphragm 1414 is from bottom to top. Since the sound heard by the user does not come from the vibration felt by the user's bones, skin, etc., but the first diaphragm 1413 and the second diaphragm 1414 change the air density by pushing the air to vibrate, so that the user can hear the sound. Therefore, without affecting the volume of the sound signal output by the air conduction speaker assembly, the intensity of the mechanical vibration (i.e., the first mechanical vibration) of the housing 1450 and the components connected to the housing 1450 (i.e., the echo signal source) can be reduced to reduce the intensity of the mechanical vibration (i.e., the third mechanical vibration) transmitted by the housing 1450 received by the bone conduction microphone (not shown in the figure), thereby reducing the intensity of the first signal generated by the bone conduction microphone. In addition, the speaker assembly 1410 is also provided with a second diaphragm 1414 having a vibration direction opposite to that of the first diaphragm 1413. Two diaphragms are provided in the air conduction speaker assembly. The mechanical vibration generated by the first diaphragm 1413 will cause the shell 1450 to vibrate, and the mechanical vibration generated by the second diaphragm 1414 will also cause the shell 1450 to vibrate. Since the vibration direction of the first diaphragm 1413 is opposite to the vibration direction of the second diaphragm 1414, the two mechanical vibrations generated on the shell cancel each other out, thereby reducing the intensity of the mechanical vibration of the shell. In some embodiments, the two diaphragms may be components in the same air conduction speaker assembly. In other embodiments, the acoustic input-output device 1400 may include a first air conduction speaker assembly and a second air conduction speaker assembly, and the first diaphragm 1413 and the second diaphragm 1414 are components in the first air conduction speaker assembly and the second air conduction speaker assembly, respectively. Fig.14 In the illustrated embodiment, it can be considered that there are two air conduction speaker assemblies, which are respectively located in different areas of the shell 1450, and each air conduction speaker assembly includes a diaphragm, a magnetic circuit assembly and a coil.

[0115] In some embodiments, the shell 1450 may include a first cavity 1455 and a second cavity 1456, and the first diaphragm 1413 and the second diaphragm 1414 may be located in the first cavity 1455 and the second cavity 1456, respectively. The shell 1450 may include a first portion corresponding to the first cavity 1455 and a second portion corresponding to the second cavity 1456. The side wall of the first cavity 1455 (i.e., the side wall of the first portion of the shell 1450) may be provided with a first sound-transmitting hole 1451 and a second sound-transmitting hole 1452. In some embodiments, the first sound-transmitting hole 1451 and the second sound-transmitting hole 1452 may be disposed on different side walls of the first portion of the shell 1450. In some embodiments, the first sound-transmitting hole 1451 and the second sound-transmitting hole 1452 may be disposed on non-adjacent side walls of the first portion of the shell 1450, i.e., the first sound-transmitting hole 1451 and the second sound-transmitting hole 1452 may be disposed on opposite sides of the first portion of the shell 1450 (e.g., Fig.14 shown).

[0116] The side wall of the second cavity 1456 (i.e., the side wall of the second part of the shell 1450) may be provided with a third sound-permeable hole 1453 and a fourth sound-permeable hole 1454. In some embodiments, the third sound-permeable hole 1453 and the fourth sound-permeable hole 1454 may be provided on different side walls of the second part of the shell 1450. In some embodiments, the third sound-permeable hole 1453 and the fourth sound-permeable hole 1454 may be provided on non-adjacent side walls of the second part of the shell 1450, that is, the third sound-permeable hole 1453 and the fourth sound-permeable hole 1454 may be provided on opposite sides of the second part of the shell 1450 (e.g., Fig.14 shown).

[0117] like Fig.14As shown, in some embodiments, the first sound-permeable hole 1451 and the third sound-permeable hole 1453 can be arranged on the same side of the housing 1450. The second sound-permeable hole 1452 and the fourth sound-permeable hole 1454 can be arranged on the same side of the housing 1450, so that the phase of the sound emitted by the first sound-permeable hole 1451 is the same as the phase of the sound emitted by the third sound-permeable hole 1453, and the phase of the sound emitted by the second sound-permeable hole 1452 is the same as the phase of the sound emitted by the fourth sound-permeable hole 1454. In this embodiment, the housing 1450 is divided into two unconnected cavities, namely the first cavity 1455 and the second cavity 1456, and the first air conduction speaker assembly or (the first vibration element 1411) and the second air conduction speaker assembly (or the second vibration element 1412) are respectively located in the two cavities. The first cavity 1455 can be divided into a front cavity and a rear cavity by the first diaphragm 1413, and the second cavity 1456 can be divided into a front cavity and a rear cavity by the second diaphragm 1414. The first sound hole 1451 and the third sound hole 1453 can be equivalent to the front sound holes of the first cavity 1455 and the second cavity 1456, and the second sound hole 1452 and the fourth sound hole 1454 can be equivalent to the back sound holes of the first cavity 1455 and the second cavity 1456. When the sound phases of the front sound holes of the first cavity 1455 and the second cavity 1456 are the same and the sound phases of the back sound holes are also the same, the sounds emitted by the two diaphragms are the same, so the volume of air conduction will not be reduced.

[0118] In some embodiments, when the speaker assembly 1410 has multiple diaphragms, the structure of the speaker assembly 1410 can be adjusted to reduce the overall size.

[0119] like Fig.15 As shown, in some embodiments, the speaker assembly 1510 may include a first vibration element 1511 and a second vibration element 1512, the first vibration element 1511 includes a first diaphragm 1513, a first magnetic circuit assembly 1515 and a first coil 1517, and similarly, the second vibration element 1512 also includes a second diaphragm 1514, a second magnetic circuit assembly 1516 and a second coil 1518 (or a voice coil), and the first cavity 1555 and the second cavity 1556 may be connected. The first magnetic circuit assembly 1515 and the second magnetic circuit assembly 1516 are connected as a whole to reduce the occupied space of the entire speaker assembly 1510.

[0120] In some embodiments, the first air conduction speaker assembly and the second air conduction speaker assembly may be two identical speakers. In some embodiments, the first air conduction speaker assembly and the second air conduction speaker assembly may be two different speakers. For example, in the acoustic input-output device 1500, a first air conduction speaker assembly and a second air conduction speaker assembly are included, wherein the first air conduction speaker assembly may be used as a main speaker, mainly generating a sound signal heard by a user. The second air conduction speaker assembly may be used as an auxiliary speaker. By adjusting the intensity of the mechanical vibration of the auxiliary speaker, it generates a force opposite to that of the main speaker on the housing 1550, thereby reducing the vibration intensity of the housing 1550. In some embodiments, the speaker assembly 1510 may include a main speaker and an auxiliary device for generating a vibration on the housing 1550 in a direction opposite to that of the main speaker. In some embodiments, the auxiliary device may be a vibration motor, and the vibration motor may generate a vibration on the housing 1550 in a direction opposite to that of the main speaker, thereby reducing the vibration intensity of the housing 1550. In some embodiments, the intensity of the mechanical vibration generated by the auxiliary speaker may be adjusted. Specifically, the speaker assembly 1510 may include an auxiliary speaker control device, which can obtain the strength and direction of the mechanical vibration of the main speaker, and adjust the strength and direction of the mechanical vibration generated by the auxiliary speaker based on the strength and direction of the mechanical vibration of the main speaker, so that the force of the auxiliary speaker on the shell and the force of the main speaker on the shell 1550 can offset each other to reduce the vibration of the shell 1550, and further reduce the vibration of the shell 1550 transmitted to the bone conduction microphone 1520 to reduce the bone conduction microphone ( Fig.15 The strength of the echo signal generated (not shown).

[0121] It should be noted that the embodiment of setting the vibration directions of the two diaphragms to be opposite can be combined with one or more of the above embodiments. For example, in the embodiment of setting the vibration directions of the two diaphragms to be opposite, a second vibration reduction structure can be set between the first diaphragm (e.g., the first diaphragm 1413) and the housing (e.g., the housing 1450) and between the second diaphragm (e.g., the second diaphragm 1414) and the housing 1450 to reduce the mechanical vibration received by the housing 1450, thereby reducing the intensity of the first mechanical vibration received by the bone conduction microphone.

[0122] In some embodiments, the voice signal source can be a vibration part when providing the voice signal to the user. For example, when the user speaks, the vibration intensity of the vocal cords, mouth, nasal cavity, throat and other parts is obviously higher than that of the ears, eyes and other parts. Therefore, these parts can be used as voice signal sources. In some embodiments, when designing the bone conduction microphone 1920, the bone conduction microphone 1920 can be located near at least one of the user's mouth, nasal cavity or vocal cords. For example, when the acoustic input and output device 1900 is Fig.19 When the glasses are worn, the bone conduction microphone 1920 can be set in the nose bridge 1935 of the glasses. Since the bone conduction microphone 1920 is close to the bridge of the user's nose, the intensity of the mechanical vibration received is greater. Fig.19 More description of the glasses shown can be found in other embodiments of the present application and will not be repeated here. Fig.19 As shown, in some embodiments, the acoustic input-output device 1900 can be set so that when the user wears the acoustic input-output device 1900, the distance between the bone conduction microphone 1920 and the vibration part of the user (not shown in the figure) is less than the third threshold. As described herein, taking the distance between the bone conduction microphone 1920 and the throat of the user as an example, in some embodiments, the third threshold may be 20 cm. In some embodiments, the third threshold may be 15 cm. In some embodiments, the third threshold may be 10 cm. In some embodiments, the third threshold may be 2 cm. In this embodiment, since the bone conduction microphone 1920 is closer to the vibration part of the user, the intensity of the received second mechanical vibration (i.e., the fourth mechanical vibration) is greater, and the greater the intensity of the second signal generated by the bone conduction microphone 1920, the stronger the voice signal intensity can be effectively improved.

[0123] Fig.16 Schematic diagram of the structure of headphones shown in some embodiments of the present application. Fig.16 As shown, in some embodiments, the acoustic input and output device 1600 can be a headset, including a fixing component 1630. The fixing component 1630 can include a headband 1632 and two earmuffs 1631 connected to both sides of the headband 1632. The headband 1632 can be used to fix the headset to the user's head and fix the two earmuffs 1631 to both sides of the user's head. A bone conduction microphone 1620 and a speaker assembly 1610 can be provided in each earmuff 1631. In some embodiments, the bone conduction microphone 1620 can be located at any position in the earmuff 1631. For example, the bone conduction microphone 1620 can be located at a position slightly above the earmuff 1631. For another example, the bone conduction microphone 1620 can be located at a position slightly below the earmuff 1631 (such as Fig.16(as shown), when the user wears the acoustic input and output device 1600, the distance between the bone conduction microphone 1620 and the vibration part of the user can be shortened. In this embodiment, the bone conduction microphone 1620 is closer to the vibration part when the user speaks, so that the vibration (i.e., the fourth mechanical vibration) intensity of the vibration part received by the bone conduction microphone 1620 when the user speaks can be greater, and the intensity of the second signal generated by the bone conduction microphone 1620 can be greater. In turn, the ratio of the intensity of the second signal to the intensity of the fourth signal is greater, the proportion of the echo signal in the sound signal generated by the bone conduction microphone is smaller, and the user experience is better.

[0124] Fig.17 Schematic diagram of the structure of a single-ear headset shown in some embodiments of the present application. Fig.17 As shown, in some embodiments, the acoustic input and output device 1700 may be a monaural headset, that is, the bone conduction microphone 1720 and the speaker assembly 1710 may be respectively disposed in two earmuffs 1731, and only one speaker assembly 1710 or one bone conduction microphone 1720 is disposed in each earmuff 1731. In this embodiment, since the bone conduction microphone 1720 and the speaker assembly 1710 are respectively disposed in different earmuffs 1731, located on both sides of the user's head, the distance between the bone conduction microphone 1720 and the speaker assembly 1710 is relatively far, so the intensity of the first mechanical vibration generated by the speaker assembly 1710 received by the bone conduction microphone 1720 is relatively small, that is, the intensity of the third mechanical vibration is relatively small, so that the proportion of the echo signal in the sound signal generated by the bone conduction microphone 1720 is relatively small, and the user experience is relatively good. In some embodiments, the headband 1732 may include one or more second vibration reduction structures (not shown in the figure) for reducing the intensity of the first mechanical vibration transmitted via the headband 1732. In some embodiments, the headband 1732 may be provided with foam to reduce the intensity of the first mechanical vibration transmitted from the speaker assembly 1710 to the bone conduction microphone 1720. In other specific embodiments, the headband 1732 may be made of a second vibration-damping material. The vibration-damping material may be the same as the vibration-damping material in one or more of the aforementioned embodiments. For example, the headband 1732 may be made of materials such as silicone or rubber.

[0125] In some embodiments, the bone conduction microphone 1720 or the speaker assembly 1710 may not be disposed in the earmuff 1731. For example, the bone conduction microphone may be disposed in the earmuff 1731. Fig.16 and Fig.17 Point D on the headband shown in the figure corresponds to the top of the user's head, and the speaker assembly is arranged in the earmuff. For another example, the speaker assembly can be arranged at Fig.16 and Fig.17 Point D on the headband is shown, which corresponds to the top of the user's head, and the bone conduction microphone is set in the ear cup.

[0126] Fig.18 2 is a cross-sectional schematic diagram of a binaural headset shown in some embodiments of the present application. Fig.16 and Fig.18 As shown, in some embodiments, the acoustic input and output device 1800 can be a binaural headset, including a fixing assembly 1830. The fixing assembly 1830 can include a headband 1832 and two earmuffs 1831 connected to both sides of the headband 1832. A sponge cover 1833 can be provided on the side of each earmuff 1831 that contacts the user's face 1840, and the bone conduction microphone 1820 can be accommodated in the sponge cover 1833. After the sponge cover 1833 is provided, it is equivalent to adding a vibration reduction structure between the bone conduction microphone 1820 and the housing 1850 of the speaker assembly 1810, that is, the second vibration reduction structure in the aforementioned embodiment, which reduces the intensity of the first mechanical vibration generated by the speaker assembly 1810 transmitted via the housing 1850. Further, since the sponge cover 1833 has a large elasticity, it will weaken the intensity of the second mechanical vibration transmitted via the user's face 1840. Therefore, in some embodiments, a part of the surface of the sponge cover 1833 can be provided with a vibration transmission structure with a large rigidity. In some embodiments, the vibration transmission structure can be set as a sheet member, for example, a metal sheet or a plastic sheet (the metal sheet and the plastic sheet are not shown in the figure). In some embodiments, the outer side of the sheet member can contact the user's face 1840, and the inner side of the sheet member is connected to the bone conduction microphone 1820. In this embodiment, the user's face 1840 is in contact with the bone conduction microphone 1820 through a sheet member with greater rigidity, so as to minimize the loss of the vibration (i.e., the second mechanical vibration) received by the bone conduction microphone 1820 during the transmission process when the user speaks, and increase the intensity of the fourth mechanical vibration, thereby increasing the intensity of the voice signal generated by the bone conduction microphone 1820.

[0127] Fig.19 Schematic diagram of the structure of glasses shown in some embodiments of the present application. Fig.19As shown, in some embodiments, the acoustic input and output device 1900 may be a pair of glasses with speaker and microphone functions. The glasses may include a fixing component, which may be a glasses frame 1930. The glasses frame 1930 may include a glasses frame 1932 and two glasses legs 1933. The glasses legs 1933 may include a temple body 1934 connected to the glasses frame 1932. At least one temple body 1934 may include a speaker assembly 1910 as described in the above embodiment of the present application. In some embodiments, the speaker assembly 1910 may include a bone conduction speaker assembly. The bone conduction speaker assembly may be located in the portion of the glasses leg 1933 that contacts the user's skin. In some embodiments, the glasses frame 1932 may include a nose bridge 1935 for supporting the glasses frame 1932 above the user's nose bridge. The nose bridge 1935 may be provided with a bone conduction microphone 1920 as described in the above embodiment of the present application. The nasal cavity is the part that vibrates when the user provides a voice signal, and its mechanical vibration intensity is relatively large. The advantage of setting the bone conduction microphone in the nose bridge 1935 is that, on the one hand, the intensity of the mechanical vibration of the voice signal received by the bone conduction microphone 1920 can be improved, and on the other hand, since the bone conduction microphone 1920 and the speaker assembly 1910 are set at different positions of the glasses, the intensity of the first mechanical vibration generated by the speaker assembly 1910 when transmitting the sound wave received by the bone conduction microphone 1920 is smaller, and the echo signal generated by the bone conduction microphone 1920 is smaller.

[0128] It should be noted that the glasses described in the above embodiments can be various types of glasses, for example, sunglasses, myopia glasses, and hyperopia glasses. In some embodiments, the glasses can also be glasses with VR (Virtual Reality) function or AR (Augmented Reality) function.

[0129] The beneficial effects that may be brought about by the embodiments of the present application include but are not limited to: (1) setting the first angle formed by the vibration direction of the bone conduction microphone and the vibration direction of the echo signal source within a set angle range, thereby reducing the vibration intensity of the echo signal source received by the bone conduction microphone and reducing the intensity of the generated echo signal (i.e., the first signal); (2) setting the second angle formed by the vibration direction of the bone conduction microphone and the vibration direction of the voice signal source within a set angle range, thereby increasing the vibration intensity of the voice signal source received by the bone conduction microphone and increasing the intensity of the generated voice signal (i.e., the second signal); (3) controlling the clamping force on the contact part of the acoustic input / output device with the user within a certain range, so that the bone conduction microphone The closer the contact with the user is, the higher the vibration intensity of the received voice signal source (i.e., the intensity of the fourth mechanical vibration); (4) a vibration reduction structure is added between the bone conduction microphone and the shell of the speaker assembly to reduce the intensity of the received vibration of the speaker assembly (i.e., the intensity of the third mechanical vibration); (5) a vibration reduction structure is added between the vibration element of the speaker assembly and the shell, and the vibration reduction structure is used to reduce the influence of the vibration of the vibration element on the shell, thereby reducing the intensity of the mechanical vibration generated by the shell, and finally reducing the intensity of the vibration of the speaker assembly received by the bone conduction microphone; (6) the bone conduction microphone is set to be closer to the vibration part when the user provides the voice signal, thereby increasing the intensity of the vibration of the received voice signal source. It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects that may be produced may be any one or a combination of the above, or any other beneficial effects that may be obtained.

[0130] The basic concepts have been described above. Obviously, for those skilled in the art, the above invention disclosure is only used as an example and does not constitute a limitation of the present application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements and amendments to the present application. Such modifications, improvements and amendments are suggested in the present application, so such modifications, improvements and amendments still belong to the spirit and scope of the exemplary embodiments of the present application.

[0131] At the same time, the present application uses specific words to describe the embodiments of the present application. For example, "one embodiment", "an embodiment" and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of the present application. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different positions in this specification does not necessarily refer to the same embodiment. In addition, some features, structures or characteristics in one or more embodiments of the present application can be appropriately combined.

[0132] In addition, unless explicitly stated in the claims, the order of the processing elements and sequences described in this application, the use of alphanumeric characters or other names, are not intended to limit the order of the processes and methods of this application. Although the above disclosure discusses some invention embodiments that are currently considered useful through various examples, it should be understood that such details are only for illustrative purposes, and the attached claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the essence and scope of the embodiments of this application. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.

[0133] Similarly, it should be noted that in order to simplify the description of the disclosure of this application and thus help understand one or more embodiments of the invention, in the above description of the embodiments of this application, multiple features are sometimes combined into one embodiment, figure or description thereof. However, this disclosure method does not mean that the features required by the object of this application are more than the features mentioned in the claims. In fact, the features of the embodiments are less than all the features of the single embodiment disclosed above.

[0134] In some embodiments, numbers describing the number of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by modifiers such as "approximately", "approximately" or "substantially" in some examples. Unless otherwise specified, "approximately", "approximately" or "substantially" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical data used in the specification and claims are approximate values, which can be changed according to the required features of individual embodiments. In some embodiments, the numerical data should consider the specified significant digits and adopt the general method of retaining the digits. Although the numerical domains and data used to confirm the breadth of the range in some embodiments of the present application are approximate values, in specific embodiments, such numerical values ​​are set as accurately as possible within the feasible range. Finally, it should be understood that the embodiments described in the present application are only used to illustrate the principles of the embodiments of the present application. Other variations may also belong to the scope of the present application. Therefore, as an example and not a limitation, the alternative configurations of the embodiments of the present application can be regarded as consistent with the teachings of the present application. Accordingly, the embodiments of the present application are not limited to the embodiments explicitly introduced and described in the present application.

Claims

1. An acoustic input and output device, include: A speaker assembly for transmitting sound waves by generating a first mechanical vibration; as well as a microphone, configured to receive a second mechanical vibration generated when a voice signal source provides a voice signal, wherein the microphone generates a first signal and a second signal respectively under the action of the first mechanical vibration and the second mechanical vibration; A first angle formed by the vibration direction of the microphone and the direction of the first mechanical vibration is within a set angle range so that within a certain frequency range, the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second signal. 2 . The acoustic input-output device according to claim 1 , wherein the first angle is within a range of 20 degrees to 90 degrees. 3 . The acoustic input-output device according to claim 2 , wherein the first angle is within a range of 75 degrees to 90 degrees. The acoustic input-output device according to claim 3 , wherein the first angle comprises 90 degrees.

5. The acoustic input-output device according to claim 1, wherein a second angle formed by the vibration direction of the microphone and the direction of the second mechanical vibration is within a set angle range so that the ratio of the intensity of the first mechanical vibration to the intensity of the first signal is greater than the ratio of the intensity of the second mechanical vibration to the intensity of the second signal. 6 . The acoustic input-output device according to claim 5 , wherein the second angle is within a range of 0 to 85 degrees. 7 . The acoustic input-output device according to claim 6 , wherein the second angle is within a range of 0 to 15 degrees.

8. The acoustic input-output device according to claim 7, wherein the second angle includes 0 degrees.

9. The acoustic input-output device according to claim 1, further comprising a vibration-damping structure, the vibration-damping structure comprising a vibration-damping material having an elastic modulus less than a first threshold, the microphone being connected to the speaker assembly via the vibration-damping structure. 10 . The acoustic input-output device according to claim 9 , wherein the vibration reduction structure has a thickness of 0.5 mm to 5 mm.

Citation Information

Patent Citations

  • Voice interaction module and voice interaction equipment

    CN111212346A

  • Earphone system and microphone device thereof

    CN112637736A