Audio playing method for controlling sound leakage and wearable sound equipment
By dividing the target audio signal into two parts, it is sent to the bone conductor speaker module and the air conductor speaker module respectively, and controlling the signal strength based on the audio source, the problem of wearable audio equipment being prone to sound leakage and insufficient sound quality when playing sound is achieved, and sound quality improvement and sound leakage control are achieved.
Patent Information
- Application Number
- CN202510181517.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-27
AI Technical Summary
Existing wearable audio equipment is prone to leak sound when playing sound, especially equipment that uses air-guided speakers. The sound has insufficient sound quality performance, making it difficult to present rich audio details and good audio effects.
By dividing the target audio signal into two parts, it is sent to the bone conductor speaker module and the air conductor speaker module respectively, and the signal intensity is controlled based on the audio source, so that the total volume of the air conductor sound and the bone conductor sound increases or decreases as the sound source volume fluctuates, and at the same time, the sound leakage of the air conductor sound module is controlled within the preset range.
It realizes sound quality improvement and sound leakage control of wearable audio equipment, ensuring reasonable sound volume and good sound quality, and effectively reducing sound leakage.
Smart Images

Figure CN120050581A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of audio playback, and particularly to an audio playback method for controlling sound leakage and a wearable audio device. Background Art
[0002] With the development of the economic society, more and more wearable audio devices that can be worn, such as AR devices or VR devices, smart glasses, smart headphones, etc., have emerged. Many wearable audio devices have the function of playing sound. When people wear wearable audio devices, they can receive various information with sound signals as the propagation carrier.
[0003] In the prior art, for wearable audio devices that use air-conduction speakers to play sound, there is a problem of easy sound leakage during sound playback. If the volume is reduced, it is not easy for users to hear clearly. For wearable audio devices that use bone-conduction speakers to play sound, there are deficiencies in the sound quality performance during sound playback, and it is difficult to present rich audio details and good audio effects.
[0004] Therefore, there is a need to provide a wearable audio device whose played sound has good sound quality, reasonable volume, and can control the sound leakage phenomenon.
[0005] The content in the background art section is only the information known to the inventor personally, and does not represent that the above information has entered the public domain before the filing date of this disclosure application, nor does it represent that it can become the prior art of this disclosure. Summary of the Invention
[0006] This specification provides an audio playback method for controlling sound leakage and a wearable audio device, which can solve the problems existing in the related art.
[0007] In a first aspect, this specification provides an audio playback method for controlling sound leakage, which is applied to a wearable audio device. The audio playback method includes: obtaining a target audio signal, where the target audio signal is a mapping of the source audio, and the source audio has fluctuations in source volume; dividing the target audio signal into a first audio signal and a second audio signal; sending the first audio signal to a bone-conduction speaker module in the audio device to emit bone-conduction sound; sending the second audio signal to an air-conduction speaker module in the audio device to emit air-conduction sound; based on the source audio, respectively controlling the signal intensities of the first audio signal and the second audio signal, so that the total volume of the air-conduction sound and the bone-conduction sound increases or decreases within a preset error following the fluctuations of the source volume, and at the same time controlling the sound leakage of the air-conduction speaker module within a preset range.
[0008] In some embodiments, where: the target audio signal includes a low-frequency signal, a mid-frequency signal, and a high-frequency signal; controlling the signal intensities of the first audio signal and the second audio signal respectively based on the source audio includes: controlling the proportions of the low-frequency signal, the mid-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal respectively based on the fluctuations of the source audio volume.
[0009] In some embodiments, where: the first audio signal includes a first low-frequency signal, a first mid-frequency signal, and a first high-frequency signal; the second audio signal includes a second low-frequency signal, a second mid-frequency signal, and a second high-frequency signal; controlling the proportions of the low-frequency signal, the mid-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal respectively based on the fluctuations of the source audio volume includes: applying a first gain and a second gain to the first audio signal and the second audio signal respectively based on the fluctuations of the source audio volume to adjust the distribution of the amplitudes of different frequencies in the first audio signal and the second audio signal, such that the proportion of the first mid-frequency signal is higher than the proportions of the first low-frequency signal and the first high-frequency signal, and the proportion of the second mid-frequency signal is lower than the proportions of the second low-frequency signal and the second high-frequency signal.
[0010] In some embodiments, the first gain and the second gain are determined based on a target leakage audio response table stored in the storage medium of the audio device.
[0011] In some embodiments, where applying the first gain and the second gain to the first audio signal and the second audio signal respectively includes: applying a first low-frequency gain, a first mid-frequency gain, and a first high-frequency gain to the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal respectively, and applying a second low-frequency gain, a second mid-frequency gain, and a second high-frequency gain to the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal respectively.
[0012] In some embodiments, where the mid-frequency signal further includes a mid-low-frequency signal and a mid-high-frequency signal; correspondingly, the first mid-frequency signal includes a first mid-low-frequency signal and a first mid-high-frequency signal, and the first gain includes a first mid-low-frequency gain and a first mid-high-frequency gain; the second mid-frequency signal includes a second mid-low-frequency signal and a second mid-high-frequency signal, and the second gain includes a second mid-low-frequency gain and a second mid-high-frequency gain; correspondingly, applying the first mid-frequency gain to the first mid-frequency signal includes: applying the first mid-low-frequency gain to the first mid-low-frequency signal and applying the first mid-high-frequency gain to the first mid-high-frequency signal; applying the second mid-frequency gain to the second mid-frequency signal includes: applying the second mid-low-frequency gain to the second mid-low-frequency signal and applying the second mid-high-frequency gain to the second mid-high-frequency signal.
[0013] In some embodiments, applying a first gain to the first audio signal and a second gain to the second audio signal respectively based on the fluctuation of the sound source volume includes: respectively controlling the values of the first low-frequency gain, the first mid-frequency gain, the first high-frequency gain, the second low-frequency gain, the second mid-frequency gain, and the second high-frequency gain based on the fluctuation of the sound source volume, so as to adjust the distribution of the first audio signal and the second audio signal in the low-frequency, mid-frequency, and high-frequency ranges based on the fluctuation of the sound source volume.
[0014] In some embodiments, the sound source audio includes low-frequency sound source audio, mid-frequency sound source audio, and high-frequency sound source audio; the low-frequency sound source audio has a low-frequency sound source volume, the mid-frequency sound source audio has a mid-frequency sound source volume, and the high-frequency sound source audio has a high-frequency sound source volume; correspondingly, the fluctuation of the sound source volume includes the fluctuation of the low-frequency sound source volume, the fluctuation of the mid-frequency sound source volume, and the fluctuation of the high-frequency sound source volume.
[0015] In some embodiments, applying a first gain to the first audio signal and a second gain to the second audio signal respectively based on the fluctuation of the sound source volume includes: when the mid-frequency sound source volume increases, reducing the second mid-frequency gain and increasing the first mid-frequency gain, and allocating a higher proportion of the energy of the mid-frequency audio to the bone conduction speaker module.
[0016] In some embodiments, applying a first gain to the first audio signal and a second gain to the second audio signal respectively based on the fluctuation of the sound source volume includes: increasing the second low-frequency gain and / or the second high-frequency gain, and reducing the first low-frequency gain and / or the first high-frequency gain, so as to keep the total volume of the air conduction sound and the bone conduction sound increasing or decreasing within a preset error following the fluctuation of the sound source volume.
[0017] In some embodiments, applying a first gain to the first audio signal and a second gain to the second audio signal respectively based on the fluctuation of the sound source volume includes: when the high-frequency sound source volume exceeds a certain threshold, reducing the second high-frequency gain and increasing the first high-frequency gain, and allocating a higher proportion of the energy of the high-frequency audio to the bone conduction speaker module.
[0018] In some embodiments, the air conduction speaker module includes at least one dipole speaker.
[0019] In some embodiments, it further includes: applying a corresponding phase difference between the first audio signal and the second audio signal based on the different distances of the bone conduction speaker module and the air conduction speaker module from the auditory nerve.
[0020] In some embodiments, controlling the proportions of the low-frequency signal, mid-frequency signal, and high-frequency signal in the first audio signal and the second audio signal respectively based on the volume fluctuation of the sound source includes: receiving a tuning instruction issued by the user indicating that the total volume is adjusted to be lower than a first preset value; based on the sound source audio, adjusting the allocation ratio of the first audio signal and the second audio signal, and reducing the proportion of the frequency signal with a higher sound intensity ratio in the first audio signal and the second audio signal compared with the sound source audio before.
[0021] In some embodiments, the sound source audio includes a low-frequency sound source audio, a mid-frequency sound source audio, and a high-frequency sound source audio; correspondingly, the first audio signal includes a first low-frequency signal, a first mid-frequency signal, and a first high-frequency signal; the second audio signal includes a second low-frequency signal, a second mid-frequency signal, and a second high-frequency signal; reducing the proportion of the frequency signal with a higher sound intensity ratio in the first audio signal and the second audio signal compared with the sound source audio before includes: based on the volume fluctuation of the sound source, adjusting the proportions of the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal in the first audio signal, and adjusting the proportions of the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal in the second audio signal.
[0022] In some embodiments, adjusting the proportions of the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal in the first audio signal includes: adjusting the first low-frequency gain, the first mid-frequency gain, and the first high-frequency gain respectively applied to the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal; adjusting the proportions of the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal in the second audio signal includes: adjusting the second low-frequency gain, the second mid-frequency gain, and the second high-frequency gain respectively applied to the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal.
[0023] In some embodiments, controlling the proportions of the low-frequency signal, mid-frequency signal, and high-frequency signal in the first audio signal and the second audio signal respectively based on the volume fluctuation of the sound source includes: receiving a tuning instruction issued by the user indicating that the total volume is adjusted to be higher than a second preset value; based on the volume fluctuation of the sound source, adjusting the allocation ratio of the first audio signal and the second audio signal so that the ratio is adjusted in the direction of maximizing the gain of the bone conduction speaker module and the air conduction speaker module in each frequency band.
[0024] In some embodiments, the audio playback method further includes increasing the preset range of the sound leakage of the speaker module.
[0025] In some embodiments, the audio playback method further includes: adjusting the allocation ratio of the first audio signal and the second audio signal based on the frequency response law of the air conduction speaker module and the frequency response law of the bone conduction speaker module, so as to adjust the allocation ratio in the direction of maximizing the gain of the bone conduction speaker module and the air conduction speaker module in each frequency band.
[0026] In some embodiments, the first audio signal includes a first low-frequency signal, a first intermediate-frequency signal, and a first high-frequency signal; the second audio signal includes a second low-frequency signal, a second intermediate-frequency signal, and a second high-frequency signal; wherein adjusting the allocation ratio of the first audio signal and the second audio signal includes: adjusting the first low-frequency gain, the first intermediate-frequency gain, and the first high-frequency gain respectively applied to the first low-frequency signal, the first intermediate-frequency signal, and the first high-frequency signal; adjusting the second low-frequency gain, the second intermediate-frequency gain, and the second high-frequency gain respectively applied to the second low-frequency signal, the second intermediate-frequency signal, and the second high-frequency signal.
[0027] In some embodiments, wherein controlling the proportions of the low-frequency signal, the intermediate-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal respectively based on the volume fluctuation of the sound source includes: receiving a tuning instruction issued by the user, indicating that the total volume is adjusted to be higher than a second preset value; increasing the preset range of the sound leakage of the speaker module; further adjusting the distribution of the amplitudes of different frequencies in the first audio signal, so that the proportion of the first intermediate-frequency signal in the first audio signal is further increased; further adjusting the distribution of the amplitudes of different frequencies in the second audio signal described, so that the proportion of the second intermediate-frequency signal is higher than the proportions of the second low-frequency signal and the second high-frequency signal.
[0028] In some embodiments, wherein controlling the proportions of the low-frequency signal, the intermediate-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal respectively based on the volume fluctuation of the sound source includes: receiving a tuning instruction issued by the user, indicating that the total volume is adjusted to be higher than a second preset value; obtaining the ambient sound of the audio device through a microphone module on the audio device; regarding the ambient sound as noise and calculating the signal-to-noise ratio of the total volume in each frequency band; adjusting the volume corresponding to the target frequency band in the first audio signal and / or the second audio signal, where the target frequency band is the frequency band with a signal-to-noise ratio higher than the preset value.
[0029] In some embodiments, the audio playback method further includes: increasing the volume of the target frequency band to a signal-to-noise ratio of not less than 10 dB.
[0030] In a second aspect, this specification provides a wearable audio device, which includes a wearing bracket, an air-conduction speaker module, a bone-conduction speaker module, and an audio processing module. The wearing bracket is configured to be wearable on the user's head; the air-conduction speaker module is disposed on the wearing bracket and faces the user's earhole and is at a preset distance from the user's earhole; the bone-conduction speaker module is disposed on the wearing bracket, behind the user's ear and close to the user's skin; the audio processing module is disposed on the wearing bracket and is communicatively connected to the air-conduction speaker module and the bone-conduction speaker module, and is configured to execute the above audio playback method during operation and send the target audio signal to the air-conduction speaker module and the bone-conduction speaker module.
[0031] In some embodiments, the wearing bracket is a spectacle frame, and the wearing bracket includes a left temple and a right temple; the air-conduction speaker module includes at least one left air-conduction speaker and at least one right air-conduction speaker; the bone-conduction speaker module includes at least one left bone-conduction speaker and at least one right bone-conduction speaker; wherein, at least one left air-conduction speaker and at least one left bone-conduction speaker are disposed on the left temple; at least one right air-conduction speaker and at least one right bone-conduction speaker are disposed on the right temple.
[0032] In some embodiments, the wearable audio device further includes a volume control module, which is disposed on the wearing bracket and communicatively connected to the audio processing module, and is configured to receive a volume control signal input by the user; the audio processing module executes the above audio playback method based on the volume control signal.
[0033] In some embodiments, the wearable audio device further includes a volume control module and a microphone module. The volume control module is disposed on the wearing bracket and communicatively connected to the audio processing module, and is configured to receive a volume control signal input by the user; the microphone module is disposed on the wearing bracket and communicatively connected to the audio processing module, and is configured to receive ambient sound around the device; the audio processing module further executes the above audio playback method based on the volume control signal.
[0034] In some embodiments, the audio processing module includes at least one storage medium and at least one processor: at least one storage medium stores at least one set of instruction sets for executing the audio playback method; at least one processor is communicatively connected to at least one storage medium, the air-conduction speaker module, the bone-conduction speaker module, and the volume control module, and executes at least one set of instruction sets during operation to execute the audio playback method, so as to send the target audio signal to the air-conduction speaker module and the bone-conduction speaker module.
[0035] As can be seen from the above technical solutions, the method and device provided in this specification can use the bone-conduction speaker module and the air-conduction speaker module to play sounds together, which is beneficial to improving the playback ability of the wearable audio device in terms of volume, so that the overall volume of the wearable audio device can still reasonably present the source audio. The working states of the bone-conduction speaker module and the air-conduction speaker module can be coordinated with each other, so as to make full use of the respective characteristics of the bone-conduction speaker module and the air-conduction speaker module to improve the sound effect of the wearable audio device playing sounds, and can also reduce the overall leakage of the wearable audio device by reducing the leakage of the air-conduction speaker module.
[0036] The audio playback method for controlling sound leakage and other functions of the wearable audio device provided in this specification will be partially listed in the following description. According to the description, the content introduced by the following numbers and examples will be obvious to those of ordinary skill in the art. The creative aspects of the audio playback method for controlling sound leakage and the wearable audio device provided in this specification can be fully explained through practice or the use of the methods, devices, and combinations described in the detailed examples below. Brief Description of the Drawings
[0037] To more clearly illustrate the technical solutions in the embodiments of this specification, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0038] Figure 1 Shows a schematic three-dimensional structure diagram of a wearable audio device provided according to some embodiments of this specification;
[0039] Figure 2 Shows a schematic block diagram of the circuit structure of a wearable audio device provided according to some embodiments of this specification;
[0040] Figure 3 Shows a flowchart of an audio playback method for controlling sound leakage provided according to some embodiments of this specification;
[0041] Figure 4 Shows a flowchart of adjusting the signal strength ratio provided according to some embodiments of this specification;
[0042] Figure 5 Shows a flowchart of adjusting the signal strength ratio provided according to some embodiments of this specification;
[0043] Figure 6 Shows a flowchart of adjusting the signal strength ratio provided according to some embodiments of this specification; and
[0044] Figure 7 Shows a flowchart of adjusting the signal strength ratio provided according to some embodiments of this specification. Detailed Description of the Embodiments
[0045] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various partial modifications to the disclosed embodiments are obvious, and the general principles defined here can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the illustrated embodiments, but has the broadest scope consistent with the claims.
[0046] The terms used herein are for the purpose of describing particular example embodiments only and are not restrictive. For example, unless the context clearly dictates otherwise, as used herein, the singular forms "a", "an", and "the" may also include the plural forms. When used in this specification, the terms "comprises", "comprising", and / or "having" mean that the associated features, integers, steps, operations, elements, and / or components exist, but do not preclude the existence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups.
[0047] In view of the following description, these features and other features of this specification, as well as the operations and functions of the related elements of the structure, and the economy of the combination and manufacture of the components can be significantly improved. Referring to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0048] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations in the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.
[0049] In this specification, the expression "X includes at least one of A, B, or C" means that X includes at least A, or X includes at least B, or X includes at least C. That is, X may include only any one of A, B, C, or may include any combination of A, B, C, AB, AC, BC, or ABC, as well as other possible contents / elements. Any combination of A, B, C may be A, B, C, AB, AC, BC, or ABC.
[0050] In this specification, unless explicitly stated, the association relationships generated between structures can be either direct or indirect. For example, when describing "A is connected to B", unless it is explicitly stated that A is directly connected to B, it should be understood that A can be directly connected to B or indirectly connected to B; for another example, when describing "A is above B", unless it is explicitly stated that A is directly above B (A and B are adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (there are other elements between A and B and A is above B). And so on.
[0051] With the development of the economic society, more and more wearable audio devices have emerged, such as AR devices, VR devices, smart glasses, smart headphones, smart helmets, head-mounted displays or the like, or any combination thereof. Many wearable audio devices have the function of playing sound. When people wear wearable audio devices, they can receive various information with sound signals as the propagation carrier through the wearable audio devices. However, due to their own structural reasons, wearable audio devices often face the following disadvantages:
[0052] First of all, for wearable audio devices that use air-conduction speakers to play sound, the sound is transmitted to the human ear through air conduction, resulting in easy leakage of sound to the external environment. When a user uses a wearable audio device in a public place, people around may hear the sound content played by the wearable audio device, resulting in the leakage of the user's privacy. When a user uses a wearable audio device in a quiet environment such as an office or a library, the sound leakage of the wearable audio device may interfere with others. If the volume of the wearable audio device is turned down, the sound of the wearable audio device does not sound clear enough. If the volume is reduced, it is not easy for the user to hear clearly. And for wearable audio devices that use bone-conduction speakers to play sound, there are deficiencies in the sound quality performance during sound playback, and it is difficult to present rich audio details and good audio effects.
[0053] Secondly, due to the volume limitation of wearable audio devices, the speaker units of wearable audio devices are often small in size. The sound of a single speaker is not full enough in the low-frequency part and may not be clear enough in the high-frequency part. In this way, the sound will sound relatively thin and dry, and it is difficult to meet the needs of users who pursue high-quality audio experiences.
[0054] Moreover, the speakers of wearable audio devices are often small in size, which also limits the sound volume of the speakers. In a relatively noisy environment, such as public places like subways and shopping malls, the sound of wearable audio devices may be masked by the surrounding environmental noise, making it difficult for users to clearly hear the sound of the wearable audio devices, thus affecting the user experience. For example, when a user uses a wearable audio device to watch a video or make a voice call outdoors, if the surrounding environmental noise is large, the sound of the wearable audio device will not sound clear enough.
[0055] Therefore, there is a need to provide a wearable audio device that plays sound with good sound quality, reasonable volume, and can control sound leakage.
[0056] In view of this, the embodiments of this specification provide an audio playback method and a wearable audio device for controlling sound leakage. The audio playback method for controlling sound leakage is applied to include: obtaining a target audio signal, where the target audio signal is a mapping of the source audio, and the source audio has fluctuations in source volume; dividing the target audio signal into a first audio signal and a second audio signal; sending the first audio signal to the bone conduction speaker module in the audio device to emit bone conduction sound; sending the second audio signal to the air conduction speaker module in the audio device to emit air conduction sound; based on the source audio, respectively controlling the signal intensities of the first audio signal and the second audio signal so that the total volume of the air conduction sound and the bone conduction sound increases or decreases following the fluctuations of the source volume within a preset error, and at the same time controlling the sound leakage of the air conduction speaker module within a preset range. In this way, the bone conduction speaker module and the air conduction speaker module can be used together to play sound, which is beneficial to improving the sound playback ability of the wearable audio device in terms of volume, so that the overall volume of the wearable audio device can still reasonably present the source audio. The working states of the bone conduction speaker module and the air conduction speaker module can be coordinated with each other, so as to make full use of the respective characteristics of the bone conduction speaker module and the air conduction speaker module to improve the sound effect of the wearable audio device playing sound, and can also reduce the overall sound leakage of the wearable audio device by reducing the sound leakage of the air conduction speaker module.
[0057] Next, the technical solutions of the embodiments of this specification will be described in detail with reference to the accompanying drawings. In this specification, a smart glasses is taken as an example to illustrate the above-mentioned wearable audio device. However, those skilled in the art can understand that other types of wearable audio devices are also applicable to the invention in this specification without departing from its spirit.
[0058] The wearable audio device may include at least multiple speakers. Compared with a single speaker, two or more speakers are beneficial to enhancing the overall volume of the wearable audio device and can also create a richer sound effect.
[0059] The plurality of speakers may include a first speaker and a second speaker. In some embodiments, the first speaker is an air conduction speaker module including one or more air conduction speakers; the second speaker is a bone conduction speaker module including one or more bone conduction speakers. For example, the wearable audio device includes a total of 2 air conduction speakers and 2 bone conduction speakers. When the user wears the wearable audio device on the head, the first speaker and the second speaker may be symmetric with respect to the midline of the user's head. In some embodiments, the air conduction speakers in the wearable audio device are non-earbud type, so when worn, the first speaker is spaced from the external auditory canal. As a bone conduction speaker, the second speaker needs to be in contact with the skin at the temporal bone to conduct sound.
[0060] At this time, when the user listens to the sound of the worn wearable audio device, compared with the bone conduction speaker, the air conduction speaker is more likely to leak sound. This is because the sound of the air conduction speaker propagates to the human ear through the air, and the sound can easily leak to the outside when passing through the air. When the sound of the bone conduction speaker is transmitted to the user, it does not pass through the external environment but is listened to by the user through the temporal bone. Therefore, the propagation path of the sound of the air conduction speaker in the external air is longer than that of the sound of the bone conduction speaker in the external air. This results in the sound of the air conduction speaker being more likely to leak to the outside than the sound of the bone conduction speaker.
[0061] However, the sound quality of the bone conduction speaker is not as good as that of the air conduction speaker. For example, the frequency response characteristic of the bone conduction speaker is not as good as that of the air conduction speaker. Among them, the frequency response characteristic, as an index to measure the sound quality, can be reflected by the frequency response curve. The frequency response curve refers to the curve of the gain of the speaker changing with frequency. An ideal frequency response curve should be flat, that is, after the audio signal (electrical signal) is converted into a sound signal by the speaker, its amplitude does not distort due to different frequencies.
[0062] Figure 1 The schematic perspective view of the wearable audio device provided according to some embodiments of the present specification is shown. Figure 2 The schematic block diagram of the circuit structure of the wearable audio device provided according to some embodiments of the present specification is shown.
[0063] Specifically, as Figure 1 and Figure 2As shown, the wearable audio device 001 includes a wearing bracket 100, an air conduction speaker module 200, a bone conduction speaker module 300, and an audio processing module 400. The wearing bracket 100 is configured to be wearable on the user's head. The air conduction speaker module 200 is disposed on the wearing bracket 100 and includes at least one air conduction speaker. The air conduction speaker faces the user's earhole or ear and is at a preset distance from the user's earhole. For example, the preset distance can be about 2 cm. The bone conduction speaker module 300 includes at least one bone conduction speaker, which is disposed on the wearing bracket 100, around the user's ear (such as behind the ear) and close to the user's skin.
[0064] In some embodiments, the wearable audio device 001 can be smart glasses, and one air conduction speaker and one bone conduction speaker can be provided on each temple of the smart glasses. Accordingly, the wearing bracket 100 is a spectacle frame, and the wearing bracket 100 includes a left temple 110 and a right temple 120; the air conduction speaker module 200 includes at least one left air conduction speaker 210 and at least one right air conduction speaker 220; the bone conduction speaker module 300 includes at least one left bone conduction speaker 310 and at least one right bone conduction speaker 320. Among them, the at least one left air conduction speaker 210 and the at least one left bone conduction speaker 310 are disposed on the left temple 110; the at least one right air conduction speaker 220 and the at least one right bone conduction speaker 320 are disposed on the right temple 120.
[0065] Further, the left temple 110 can include a connecting portion 111 and a fixing portion 112. The connecting portion 111 is in a straight rod shape, and its two ends are respectively connected to the spectacle frame and the fixing portion 112. The left temple 110 is bent at the connection between the connecting portion 111 and the fixing portion 112, so that when the user wears the wearable audio device 001, the fixing portion 112 can hook the left ear from behind the user's left ear, thereby fixing the smart glasses on the user's head.
[0066] The structure of the right temple 120 is the same as or similar to that of the left temple 110. The right temple 120 can include a connecting portion and a fixing portion. The connecting portion is in a straight rod shape, and its two ends are respectively connected to the spectacle frame and the fixing portion. The right temple 120 is bent at the connection between the connecting portion and the fixing portion, so that when the user wears the wearable audio device 001, the fixing portion can hook the right ear from behind the user's right ear, thereby fixing the smart glasses on the user's head
[0067] The left air-conduction speaker 210 can be disposed at the connecting portion 111 of the left spectacle temple 110, and the left bone-conduction speaker 310 is disposed at the fixing portion 112 of the left spectacle temple 110 or at the connection between its connecting portion 111 and the fixing portion 112. In this way, when the user wears the wearable audio device 001, the left air-conduction speaker 210 is located in front of the user's left ear near the cheek and faces the user's left ear; the left bone-conduction speaker 310 is located above or behind the user's left ear and contacts the skin at the temporal bone of the human body.
[0068] Similarly, the right air-conduction speaker 220 can be disposed at the connecting portion 111 of the right spectacle temple 120, and the right bone-conduction speaker 310 is disposed at the fixing portion of the right spectacle temple 120 or at the connection between its connecting portion and the fixing portion. In this way, when the user wears the wearable audio device 001, the right air-conduction speaker 210 is located in front of the user's right ear near the cheek and faces the user's right ear; the right bone-conduction speaker 310 is located above or behind the user's right ear and contacts the skin at the temporal bone of the human body.
[0069] The audio processing module 400 is disposed on the wearing bracket 100 and is communicatively connected to the air-conduction speaker module 200 and the bone-conduction speaker module 300, and is configured to execute the audio playback method described in this specification during operation to send the target audio signal to the air-conduction speaker module 200 and the bone-conduction speaker module 300.
[0070] In some embodiments, the wearable audio device 001 further includes a volume control module 500. The volume control module 500 is disposed on the wearing bracket 100 and is communicatively connected to the audio processing module 400, and is configured to receive a volume control signal input by the user; the audio processing module 400 executes the audio playback method described in this application based on the volume control signal.
[0071] Wherein, both the audio processing module 400 and the volume control module 500 are hardware circuits. The volume control module 500 may include a volume button 501. The user can apply a physical force to the volume button 501 to adjust the playback volume of the wearable audio device 001, that is, when the audio processing module 400 receives a tuning instruction generated by the user triggering the volume button 501, it can adjust the gain of the speaker for the target audio signal to change the volume.
[0072] In some embodiments, the wearable audio device 001 further includes a volume control module 500 and a microphone module 600. The volume control module 500 is disposed on the wearing bracket 100 and communicatively connected to the audio processing module 400, and is configured to receive a volume control signal input by a user; the microphone module 600 is disposed on the wearing bracket 100 and communicatively connected to the audio processing module 400, and is configured to receive ambient sound around the device; the audio processing module 400 further executes the following audio playback method based on the volume control signal.
[0073] In some embodiments, the audio processing module 400 includes at least one storage medium 420 and at least one processor 410: the at least one storage medium 420 stores at least one set of instruction sets for playing audio; the at least one processor 410 is communicatively connected to the at least one storage medium 420, the air conduction speaker module 200, the bone conduction speaker module 300, and the volume control module 500, and executes the at least one set of instruction sets during operation to execute an audio playback method, so as to send a target audio signal to the air conduction speaker module 200 and the bone conduction speaker module 300.
[0074] The storage medium 420 may include one or more of a magnetic disk, a read-only storage medium, or a random access storage medium. The storage medium 420 may further include a non-volatile random access memory.
[0075] The processor 410 may be in the form of one or more processors 410. According to some embodiments of the present specification, the processor 410 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a microcontroller unit (MCU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof.
[0076] For purposes of illustration only, only one processor 410 is described in the wearable audio device 001 in this specification. However, it should be noted that the wearable audio device 001 in this specification may also include multiple processors 410, and the number of processors 410 included in the wearable audio device 001 is not limited in this specification. Therefore, the operations and / or method steps disclosed in this specification can be executed by one processor 410 as described in this specification, or jointly executed by multiple processors 410. For example, if the processor 410 of the wearable audio device 001 executes step A and step B in this specification, it should be understood that step A and step B can also be jointly or separately executed by two different processors 410 (for example, step A is executed by the first processor, step B is executed by the second processor, or step A and B are jointly executed by the first and second processors).
[0077] In some embodiments, the wearable audio device 001 may have a mobile payment function.
[0078] In some embodiments, the wearable audio device 001 may include a camera 130. The camera 130 can take a photo of or scan a payment code. The processor 410 can obtain a photo of the payment code through the camera 130 and extract the information of the payment code based on the photo of the payment code. In this way, the user can perform mobile payment through the wearable audio device 001. Of course, the camera 130 can also pick up other images other than the payment code.
[0079] In some embodiments, the wearable audio device 001 can communicate and connect with a cloud server. The wearable audio device 001 can confirm the payment chain through the cloud server.
[0080] In some embodiments, the wearable audio device 001 can communicate and connect with a mobile terminal device such as a mobile phone. The wearable audio device 001 can send the information of the payment code to the mobile terminal device, and the user can input information such as the payment amount and password through the mobile terminal device. In other embodiments, after the wearable audio device 001 obtains the information of the payment code, it waits for the user to confirm whether to pay. After the user confirms the payment, the wearable audio device 001 can perform mobile payment.
[0081] Furthermore, the wearable audio device 001 includes spectacle lenses 140. The photographing direction of the camera 130 is the same as the direction in which the spectacle lenses 140 face the outside world. When the user wears the wearable audio device 001 and faces the payment code directly, the camera can be aligned with the payment code.
[0082] In some embodiments, the camera includes a photographing button and / or a blink detection module, and the camera is triggered by the photographing button and / or the blink detection module to start photographing.
[0083] In some embodiments, the wearable audio device 001 includes a Bluetooth module. The wearable audio device 001 can interact with a remote payment system via the Bluetooth module for mobile payment.
[0084] In some embodiments, the wearable audio device 001 includes a display screen. The display screen can display a payment code for an external camera to capture.
[0085] The above is a schematic solution of the wearable audio device 001 in the embodiments of this specification. It should be noted that the schematic solution of the wearable audio device 001 and the technical solution of the following audio playback method belong to the same concept. The detailed content in the schematic solution of the wearable audio device 001 and the technical solution of the audio playback method can be referred to each other.
[0086] Figure 3 The flowchart of an audio playback method for controlling sound leakage provided according to some embodiments of this specification is shown. This audio playback method for controlling sound leakage P100 is applied to the wearable audio device 001. Specifically, this audio playback method can be executed by the audio processing module 400. Specifically, the processor 410 is communicatively connected to the storage medium 420, the air conduction speaker module 200, the bone conduction speaker module 300, and the volume control module 500, and executes the process steps in the audio playback method P100 by reading the instruction set in the storage medium 420. For ease of description, the execution subject of the audio playback method will be collectively referred to as the audio processing module 400 hereinafter. This method includes the following steps:
[0087] S100: Obtain a target audio signal, where the target audio signal is a mapping of the source audio, and the source audio has fluctuations in source volume.
[0088] In some embodiments, the target sound source refers to the original source of the sound, such as a concert venue, a singer singing in a recording studio, the wind and bird sounds in nature, etc. The sound emitted by the target sound source is the source audio. Generally, this source audio has fluctuations in volume. The source audio is converted into an electrical signal by a sound collection device (such as a microphone), which is called a source audio signal. The source audio signal records the sound information from the source audio. Thereafter, the source audio signal is either retained as it is or becomes a target audio signal after post-processing (such as noise reduction, sharpening, changing the sound quality, etc.). Therefore, the target audio signal is a mapping of the source audio. In some embodiments, this target audio signal is input to the audio processing module 400 of the wearable audio device 001. At this time, the target audio signal becomes the source signal of the bone conduction speaker module 300 and the air conduction speaker module 200.
[0089] Next, the audio processing module 400 divides the target audio signal according to the frequency response characteristics of the bone conduction speaker module 300 and the air conduction speaker module 200, and inputs the target audio signals in the corresponding frequency ranges into the bone conduction speaker module 300 and the air conduction speaker module 200 respectively. The specific steps are as follows.
[0090] S200: Divide the target audio signal into a first audio signal and a second audio signal.
[0091] In this step, the audio processing module 400 can divide the target audio signal according to the frequency range.
[0092] According to the human ear's perception of the sound played by the speaker, the sound frequency band can be divided into a low-frequency band, a mid-frequency band, and a high-frequency band. For example, according to some frequency definition methods, the frequency range corresponding to the low-frequency band sound is below 1KHz, the frequency range corresponding to the mid-frequency band sound is 1KHz - 10KHz, and the frequency range corresponding to the high-frequency band sound is above 10KHz. Correspondingly, the target audio signal also includes a low-frequency signal, a mid-frequency signal, and a high-frequency signal. The low-frequency signal is in the low-frequency band; the mid-frequency signal is in the mid-frequency band; the high-frequency signal is in the high-frequency band.
[0093] The audio processing module 400 divides the target audio signal into two parts, namely a first audio signal and a second audio signal. In some embodiments, the wearable audio device 001 includes a frequency detection module. The frequency detection module can detect the frequency composition of the target audio signal in real time. In some embodiments, the wearable audio device 001 includes a frequency separation module, and the frequency separation module is used to separate the low-frequency band, mid-frequency band, and high-frequency band in the target audio signal from each other. The frequency of the target audio signal can be divided into three types: low-frequency band, mid-frequency band, and high-frequency band. For example, the frequency of the target audio signal can include 50Hz and 1500Hz, 50Hz belongs to the low-frequency band, and 1500Hz belongs to the mid-frequency band. The audio processing module 400 can obtain the frequency composition of the target audio through the frequency detection module. Correspondingly, the audio processing module 400 divides the low-frequency signal in the target audio signal into a first low-frequency signal and a second low-frequency signal; divides the mid-frequency signal in the target audio signal into a first mid-frequency signal and a second mid-frequency signal; divides the high-frequency signal in the target audio signal into a first high-frequency signal and a second high-frequency signal.
[0094] Furthermore, the mid-frequency band can be divided into a mid-low frequency band and a mid-high frequency band. For example, according to some frequency definition methods, the sound frequency range corresponding to the mid-low frequency band is 1KHz to 5KHz, and the sound frequency range corresponding to the mid-high frequency band is 5KHz to 10KHz. Correspondingly, the intermediate frequency signal also includes a mid-low frequency signal and a mid-high frequency signal, where the mid-low frequency signal is located in the mid-low frequency band and the mid-high frequency signal is located in the mid-high frequency band. Since the audio processing module 400 divides the target audio signal into a first audio signal and a second audio signal, accordingly, the first intermediate frequency signal includes a first mid-low frequency signal and a first mid-high frequency signal; the second intermediate frequency signal includes a second mid-low frequency signal and a second mid-high frequency signal.
[0095] S300: Send the first audio signal to the bone conduction speaker module 300 in the audio device to emit bone conduction sound; send the second audio signal to the air conduction speaker module 200 in the audio device to emit air conduction sound.
[0096] The bone conduction speaker module 300 and the air conduction speaker module 200 play sound together, which is beneficial to improving the playback ability of the wearable audio device 001 in terms of volume, so that the overall volume of the wearable audio device 001 can still reasonably present the source audio.
[0097] On the basis of the above steps, the audio processing module 400 also makes full use of the respective characteristics of the bone conduction speaker module 300 and the air conduction speaker module 200, and by adjusting the signal input of the first audio signal and the second audio signal in different frequency ranges, the bone conduction speaker module 300 and the air conduction speaker module 200 respectively output their respective superior audio, thereby improving the overall playback sound effect of the wearable audio device 001; at the same time, according to the situation of different environments, by controlling the sound leakage of the air conduction speaker module 200, the overall sound leakage of the wearable audio device 001 is controlled.
[0098] Therefore, different environmental situations will cause the audio processing module 400 to adopt different strategies to control the first audio signal and the second audio signal. Specifically:
[0099] S400: Based on the source audio, respectively control the signal intensities of the first audio signal and the second audio signal, so that the total volume of the air conduction sound and the bone conduction sound increases or decreases following the fluctuations of the source volume within a preset error, and at the same time control the sound leakage of the air conduction speaker module 200 within a preset range.
[0100] Specifically, S400 may further include: respectively controlling the proportions of low-frequency signals, mid-frequency signals, and high-frequency signals in the first audio signal and the second audio signal based on the fluctuations in the volume of the sound source and specific playback requirements. For example, applying a first gain and a second gain to the first audio signal and the second audio signal respectively to adjust the distribution of the amplitudes corresponding to different frequencies in the first audio signal and the second audio signal.
[0101] The volume allocation ratio between the air-conduction speaker module 200 and the bone-conduction speaker module 300 is related to the frequency composition of the target audio signal.
[0102] As described above, the wearable audio device 001 includes a frequency detection module. The frequency detection module can detect the frequency composition of the target audio signal in real time. The audio processing module 400 can obtain the frequency composition of the target audio signal through the frequency detection module, and then adopt a gain strategy corresponding to the frequency composition of the target audio signal according to actual needs, and control the air-conduction speaker module 200 and the bone-conduction speaker module 300 to perform the playback work based on this gain strategy.
[0103] Furthermore, the audio processing module 400 can formulate the volume allocation ratio of the air-conduction speaker module 200 and the bone-conduction speaker module 300 at different frequencies based on the frequency composition and volume distribution of the target audio signal, that is, formulate the signal intensity distribution of the first audio signal and the second audio signal at different frequencies. This is equivalent to formulating the gain of the first audio signal and the second audio signal at different frequencies when the target audio signal is divided into the first audio signal and the second audio signal, so as to control the air-conduction speaker module 200 and the bone-conduction speaker module 300 to perform playback according to this intensity distribution.
[0104] Therefore, the control of the signal intensities of the first audio signal and the second audio signal respectively based on the source audio in step S400 means controlling the proportions of low-frequency signals, mid-frequency signals, and high-frequency signals in the first audio signal and the second audio signal respectively based on the fluctuations in the volume of the sound source and specific playback requirements, that is, respectively controlling the amplitude distributions of low-frequency signals, mid-frequency signals, and high-frequency signals in the first audio signal and the second audio signal. Among them, the distribution of the amplitudes of different frequencies in the first audio signal and the second audio signal can be reflected as: the distribution of different energies or powers in the first audio signal and the second audio signal, or can also be reflected as the allocation ratio of the gains of the speaker to the first audio signal and the second audio signal.
[0105] As described above, the audio processing module 400 can formulate the volume allocation ratio of the air-conduction speaker module 200 and the bone-conduction speaker module 300 at different frequencies based on the frequency composition and volume distribution of the target audio signal. For example:
[0106] (A) When the user is in a quiet environment, there is no interference from ambient noise, and improving the sound quality played by the wearable audio device 001 and reducing sound leakage will be the key points of playback. Therefore, the processing method of the audio processing module 400 will focus on reducing the sound leakage of the air conduction speaker module 200 and balancing the overall frequency response characteristics of the air conduction speaker module 200 and the bone conduction speaker module 300 to improve the sound quality.
[0107] (B) When the user is using the wearable audio device 001 and adjusts the volume from a high volume to a medium volume or a lower volume, it means that the user may be in a quiet environment or has entered a quiet environment from a noisy environment. At this time, the key points of playback should be controlling the sound quality and sound leakage.
[0108] (C) When the user is using the wearable audio device 001 and adjusts the volume from a low volume or a medium volume to a high volume, it indicates that the user is in a noisy environment and may not be able to hear clearly. At this time, the key point of playback should be to give priority to increasing the volume to enhance the recognizability of the played sound, and controlling the sound quality and sound leakage are secondary. Therefore, the processing method of the audio processing module 400 will focus on enhancing the sound playback intensity in the frequency range that the human ear is most sensitive to.
[0109] Therefore, different environmental conditions will cause the audio processing module 400 to adopt different strategies to control the first audio signal and the second audio signal. Next, this application will introduce the different processing methods adopted by the audio processing module 400 for the first audio signal and the second audio signal respectively for the above three situations.
[0110] When the user is in a quiet environment, there is no interference from ambient noise, and improving the sound quality played by the wearable audio device 001 and reducing sound leakage will be the key points of playback.
[0111] On the one hand, in the case of playing low-frequency sounds, the sound quality of the air conduction speaker is better than that of the bone conduction speaker. Therefore, when playing low-frequency sounds, the gain and / or power distribution ratio of the bone conduction speaker is lower than that of the air conduction speaker. Specifically, when the sound switches from the mid-frequency band to the low-frequency band, the audio processing module 400 allocates more energy to the air conduction speaker, and the energy allocated to the bone conduction speaker decreases. Correspondingly, the volume of the air conduction speaker increases, and the volume of the bone conduction speaker decreases.
[0112] On the other hand, the directionality of low-frequency sounds is not obvious, that is, it is not easy for the human ear to identify the direction of low-frequency sounds. When there is a volume transfer in the low-frequency band between the air-conduction speaker and the bone-conduction speaker, in terms of the user's listening experience, the sound direction conversion caused by the volume transfer in the low-frequency band is not obvious. Correspondingly, it is not easy for the user to perceive the conversion of the sound direction. Therefore, the volume distribution ratio between the air-conduction speaker and the bone-conduction speaker can be quickly adjusted, and the sound quality can be quickly improved on the premise of maintaining the sound stability in the user's listening experience.
[0113] Based on similar considerations, when playing mid-frequency sounds, both the air-conduction speaker and the bone-conduction speaker have good sound quality, but the sound leakage phenomenon of the air-conduction speaker is relatively obvious. In addition, the human ear is also more sensitive to mid-frequency sounds. Therefore, when playing mid-frequency sounds, the audio processing module 400 allocates more energy to the bone-conduction speaker for playback, that is, the gain and / or power distribution ratio of the bone-conduction speaker is higher than that of the air-conduction speaker. Correspondingly, the volume of the bone-conduction speaker increases, and the volume of the air-conduction speaker decreases.
[0114] At the same time, the directionality of mid-frequency sounds is obvious, that is, it is easy for the human ear to identify the direction of mid-frequency sounds. When there is a volume transfer in the mid-frequency band between the air-conduction speaker and the bone-conduction speaker, in terms of the user's listening experience, the sound direction conversion caused by the volume transfer in the mid-frequency band is relatively obvious. Correspondingly, it is easy for the user to perceive the conversion of the sound direction. Therefore, the volume distribution ratio between the air-conduction speaker and the bone-conduction speaker needs to be adjusted slowly, which is beneficial to maintaining the sound stability in the user's listening experience. Therefore, in some embodiments, the volume distribution ratio between the air-conduction speaker and the bone-conduction speaker in the low-frequency band is adjusted at a first speed, and the volume distribution ratio between the air-conduction speaker and the bone-conduction speaker in the mid-frequency band is adjusted at a second speed, and the second speed is greater than the first speed.
[0115] Based on similar considerations, when playing high-frequency sounds, the sound quality of the air-conduction speaker is better than that of the bone-conduction speaker, and the sound leakage of the air-conduction speaker is relatively small. Therefore, when playing high-frequency sounds, the audio processing module 400 can allocate more sound to the air-conduction speaker, and the gain and / or power distribution ratio of the bone-conduction speaker is lower than that of the air-conduction speaker.
[0116] Therefore, when the user is in a quiet environment, step S400 can be further described as: Based on the volume fluctuation of the sound source, apply a first gain and a second gain to the first audio signal and the second audio signal respectively, so as to adjust the distribution of different frequency amplitudes in the first audio signal and the second audio signal, such that the proportion of the first intermediate-frequency signal is higher than the proportion of the first low-frequency signal and the proportion of the first high-frequency signal, and the proportion of the second intermediate-frequency signal is lower than the proportion of the second low-frequency signal and the proportion of the second high-frequency signal. Further, the proportion of the first intermediate-frequency signal is higher than the sum of the proportion of the first low-frequency signal and the proportion of the first high-frequency signal, and the proportion of the second intermediate-frequency signal is lower than the sum of the proportion of the second low-frequency signal and the proportion of the second high-frequency signal.
[0117] For example, in some embodiments, the audio processing module 400 can apply a first low-frequency gain, a first intermediate-frequency gain, and a first high-frequency gain to the first low-frequency signal, the first intermediate-frequency signal, and the first high-frequency signal respectively; and the audio processing module 400 can apply a second low-frequency gain, a second intermediate-frequency gain, and a second high-frequency gain to the second low-frequency signal, the second intermediate-frequency signal, and the second high-frequency signal respectively.
[0118] In some embodiments, the audio processing module 400 can adopt a more delicate gain adjustment means. For example, since the intermediate-frequency signal includes a mid-low-frequency signal and a mid-high-frequency signal, the first gain can further include a first mid-low-frequency gain and a first mid-high-frequency gain, and the audio processing module 400 applies the first mid-low-frequency gain to the first mid-low-frequency signal and the first mid-high-frequency gain to the first mid-high-frequency signal. At the same time, the second gain can further include a second mid-low-frequency gain and a second mid-high-frequency gain, and the audio processing module 400 applies the second mid-low-frequency gain to the second mid-low-frequency signal and the second mid-high-frequency gain to the second mid-high-frequency signal.
[0119] Correspondingly, the audio processing module 400 can control the values of the first low-frequency gain, the first intermediate-frequency gain, the first high-frequency gain, the second low-frequency gain, the second intermediate-frequency gain, and the second high-frequency gain respectively based on the volume fluctuation of the sound source, so as to adjust the distribution of the first audio signal and the second audio signal in the low frequency, intermediate frequency, and high frequency based on the volume fluctuation of the sound source.
[0120] The above method of performing different gain processing on target audio at different frequencies does not consider the change in the volume of the target audio signal itself. In reality, since the sound quality and volume of the sound source itself change over time, the corresponding sound source audio also fluctuates in volume over time, and this change in volume is a function of both time and frequency. That is to say, the volume change of audio signals with different frequencies in the sound source audio is different over time. Specifically, if the sound source audio is described according to frequency ranges, the sound source audio includes low-frequency sound source audio, mid-frequency sound source audio, and high-frequency sound source audio; the low-frequency sound source audio has a low-frequency sound source volume, the mid-frequency sound source audio has a mid-frequency sound source volume, and the high-frequency sound source audio has a high-frequency sound source volume. Correspondingly, the sound source volume fluctuation includes low-frequency sound source volume fluctuation, mid-frequency sound source volume fluctuation, and high-frequency sound source volume fluctuation.
[0121] Therefore, if the change in the volume of the sound source audio itself is considered, the above-mentioned applying the first gain and the second gain to the first audio signal and the second audio signal respectively may further include the following steps:
[0122] When the mid-frequency sound source volume increases, the audio processing module 400 reduces the second mid-frequency gain and increases the first mid-frequency gain, and allocates a higher proportion of the energy of the mid-frequency audio to the bone conduction speaker module 300. At the same time, the audio processing module 400 increases the second low-frequency gain and / or the second high-frequency gain, and reduces the first low-frequency gain and / or the first high-frequency gain to keep the total volume of the air conduction sound and the bone conduction sound within a preset error and increase or decrease following the fluctuation of the sound source volume.
[0123] When, following the fluctuation of the sound source volume, the high-frequency sound source volume exceeds a certain threshold, at this time, the audio processing module 400 reduces the second high-frequency gain and increases the first high-frequency gain, and allocates a higher proportion of the energy of the high-frequency audio to the bone conduction speaker module 300.
[0124] The reason for doing this is that when the waveform of the target audio signal is clipped, it indicates that the sound volume of the target audio signal is relatively large and there will be distortion. The audio processing module 400 can determine that the sound of the target audio signal is distorted based on the clipping phenomenon of the target audio signal, and can allocate more gain / volume / power to the bone conduction speaker module 300, then the gain and / or power allocation ratio of the bone conduction speaker module 300 is higher than that of the air conduction speaker module 200. Or, when the volume corresponding to certain time windows in the high-frequency band exceeds a preset threshold, there will be sound leakage in the air conduction speaker module 200. When the audio processing module 400 determines that the volume corresponding to certain time windows in the high-frequency band exceeds the preset threshold, when the target audio signal is played to the corresponding time window, it can allocate more gain / volume / power to the bone conduction speaker module 300, then the gain and / or power allocation ratio of the bone conduction speaker module 300 is higher than that of the air conduction speaker module 200.
[0125] In some embodiments, the first gain and the second gain are determined based on a target leakage response table stored in the storage medium 420 of the audio device.
[0126] To this end, the technician will first analyze the sound leakage of the wearable audio device 001, obtain a target sound leakage response table, and then store the target sound leakage response table in the storage medium of the wearable audio device 001. In this way, when playing audio, the audio processing module 400 will determine the sound frequency band that is prone to leakage according to the target sound leakage response table, and most of the volume / power of the sound frequency band that is prone to leakage will be played by the air conduction speaker module 200, thereby reducing sound leakage.
[0127] The method of obtaining the target leakage audio response table can refer to the following content.
[0128] The technician can use the sound receiving component to detect the sound leakage of the wearable audio device 001. The sound receiving component may include at least one microphone.
[0129] When analyzing the sound leakage of the wearable audio device 001, the wearable audio device 001 is located at the central detection position, and at least one microphone collects sound at different positions around the central detection position. The sound collection of the microphone can reflect the sound leakage of the wearable audio device 001. Based on the sound collection data of the microphone, the detector and / or the audio processing module 400 obtains a target sound leakage response table.
[0130] In some embodiments, the microphone of the sound receiving assembly is located in different directions of the central detection position to detect the sound leakage of the speaker of the sound receiving assembly in a specific direction. Further, a microphone is set in at least one of the directions such as the top, front top, upper left, upper right, front, rear, left, right, left front, right front, left rear, and right rear. For example, microphones are set in the eight directions of the front, rear, left, right, front left, front right, rear left, and rear right of the central detection position.
[0131] In some embodiments, the distance between the microphone and the central detection position may be 0.05m to 1m. For example, the distance between the microphone and the central detection position may be 0.06m, 0.1m, 0.2m, 0.3m or 0.8m. Furthermore, the distance between each microphone and the central detection position is 0.5m.
[0132] In some embodiments, an artificial head is placed at the central detection position, and the artificial head is used to simulate the user's head. The wearable audio device 001 can be worn by the artificial head. The artificial head can reflect and absorb acoustic wave signals like the user's head, so that the sound leakage detection environment of the wearable audio device 001 can be closer to the actual working environment of the wearable audio device 001, making the sound leakage situation of the wearable audio device 001 in the detection environment more consistent with that in the actual working environment, and improving the accuracy of the sound leakage detection data.
[0133] In some embodiments, all the speakers of the wearable audio device 001 play sounds simultaneously, and all the microphones of the sound collection component are in the on state during this sound playback process. The number of microphones of the sound collection component is m, and all the microphones of the sound collection component obtain m sound leakage audio data in total. In this way, the sound leakage situation when all the speakers play sounds simultaneously can be obtained, reducing the influence of the acoustic wave interference of different speakers on the detection accuracy of the sound leakage situation.
[0134] In some other embodiments, all the speakers of the wearable audio device 001 play sounds in sequence, and all the microphones of the sound collection component are in the on state during this sound playback process. Each microphone of the sound collection component can collect the sound of each speaker of the wearable audio device 001 to obtain the sound leakage audio data corresponding to the speaker. In this way, the microphone can obtain the individual sound leakage situation of each speaker.
[0135] The number of speakers of the wearable audio device 001 is n, and each microphone of the sound collection component can obtain n sound leakage audio data. For example, the wearable audio device 001 includes a total of 2 air conduction speakers and 2 bone conduction speakers, and each microphone of the sound collection component can obtain 4 sound leakage audio data.
[0136] The number of microphones of the sound collection component is m, and all the microphones of the sound collection component obtain n×m sound leakage audio data in total. For example, m = 8, and the 8 microphones are respectively arranged in 8 directions: directly in front of, directly behind, directly to the left of, directly to the right of, front left, front right, rear left, and rear right at the central detection position.
[0137] In some embodiments, when testing the sound leakage situation, the sound source of the speaker is white noise. White noise is noise whose power spectral density is constant throughout the frequency domain and has the same energy density at all frequencies.
[0138] In some embodiments, a technician can convert each sound leakage audio data from the time domain form to the frequency domain form through Fourier transform (FFT), then map the frequency of the sound leakage audio data to multiple sub-bands, and separately calculate the band energy of each sound leakage audio data in each sub-band. Among them, the sound leakage audio data in the frequency domain form can be a frequency response curve. For example, the frequency of the sound leakage audio data can be mapped to 64 sub-bands.
[0139] Technicians can select multiple frequency points of the leaking audio data for Fourier transform (FFT). The step size between adjacent frequency points is equal. For example, the step size between adjacent frequency points is 31.25 Hz. The number of frequency points is greater than the number of sub-bands. Dividing multiple frequency points into a smaller number of sub-bands can reduce the amount of calculation. For example, 512 frequency points can be divided into 64 sub-bands.
[0140] The frequency interval widths of different sub-bands can be the same or different. In some embodiments, the number of frequency points within each sub-band divided into the low-frequency band, the middle-frequency band, and the high-frequency band is the same. In some embodiments, since the human ear is more sensitive to the sound in the middle-frequency band, the number of frequency points within each sub-band divided into the middle-frequency band is greater than the number of frequency points within each sub-band divided into the low-frequency band or the high-frequency band. In other embodiments, since the sound in the low-frequency band is more concerned, the number of frequency points within each sub-band divided into the low-frequency band is greater than the number of frequency points within each sub-band divided into the middle-frequency band or the high-frequency band.
[0141] The number of sub-bands is denoted as p, the number of leaking audio data is denoted as n×m, and there are a total of n×m×p sub-bands.
[0142] Each microphone can obtain n leaking audio data, and the n leaking audio data can correspond to n speakers one by one. The sum of the band energies of the n leaking audio data can reflect the leaking situation of the n speakers when playing sound simultaneously in this direction and / or sub-band.
[0143] The leaking audio data collected by the microphone in each direction represents the actual leaking situation of the wearable audio device 001 in this direction. The volume allocation ratio between the air-conduction speaker and the bone-conduction speaker can be determined based on the actual leaking situation in different directions and / or sub-bands.
[0144] In some embodiments, in different directions and / or sub-bands, the leaking volume of the wearable audio device 001 is different. In different application scenarios, the leaking control levels of the wearable audio device 001 in different directions and / or sub-bands are also different. The leaking control level can include a first level and a second level. The leaking control corresponding to the first level is more stringent than the leaking control corresponding to the second level, that is, the allowed leaking volume under the first level is less than the allowed leaking volume under the second level.
[0145] For the directions and / or sub-bands with relatively strict leaking control, the leaking control level is the first level. Most of the volume / power of the sound that is likely to leak in this direction and / or sub-band can be played by the bone-conduction speaker, thereby reducing leakage.
[0146] Further, a technician can determine the leakage control levels of the wearable audio device 001 in different orientations and / or sub-bands according to the application scenarios of the wearable audio device 001. Among them, the technician can be a user or the audio processing module 400. For example, the technician can obtain the personnel distribution in the environment. The leakage control is relatively strict in the direction where there is personnel distribution, and the leakage control level is the first level. The leakage control is relatively loose in the direction where there is no personnel distribution, and the leakage control level is the second level. Also for example, the technician can obtain the noise in the environment. When the noise intensity is lower than the preset decibel threshold, the leakage control is relatively strict, and the leakage control level is the first level. When the noise intensity is higher than the preset decibel threshold, the leakage control is relatively loose, and the leakage control level is the second level. The preset decibel threshold can be 20 dB, 30 dB, 40 dB, 50 dB or 60 dB.
[0147] When determining the volume allocation ratio between the air conduction speaker and the bone conduction speaker, the actual leakage situation in each orientation and / or sub-band has different weight coefficients, and the weight coefficients are related to the leakage control levels. Among them, the volume allocation between the air conduction speaker and the bone conduction speaker can be reflected by setting different gains for the target audio signal.
[0148] For example, for the orientation and / or sub-band with relatively strict leakage control, the actual leakage situation in this orientation and / or sub-band has a higher weight coefficient. Most of the volume / power of the target audio signal corresponding to this orientation and / or sub-band can be played by the bone conduction speaker, that is, the gain of the bone conduction speaker for the target audio signal is greater than that of the air conduction speaker. For the orientation and / or sub-band with relatively loose leakage control, the actual leakage situation in this orientation and / or sub-band has a lower weight coefficient. Most of the volume / power of the target audio signal corresponding to this orientation and / or sub-band can be played by the air conduction speaker, that is, the gain of the bone conduction speaker for the target audio signal is less than that of the air conduction speaker.
[0149] In some embodiments, when the leakage control levels of all orientations are the same, the weight coefficients corresponding to all orientations can be the same, and the audio processing module 400 can determine the gain for the target audio signal based on the average value of the leakage volumes of all orientations. In other embodiments, the weight coefficients corresponding to all orientations can be the same, and the audio processing module 400 can determine the gain for the target audio signal based on the leakage volume of the orientation with the strictest leakage control.
[0150] Based on data such as leakage audio data, leakage control levels, weight coefficients, volume allocation ratios, and gains for the target audio signal, a technician can formulate target frequency response data. The set of target frequency response data is the target leakage audio frequency response table. When the wearable audio device 001 plays sound, the air conduction speaker and the bone conduction speaker can perform the sound playback work according to the target frequency response data in the target leakage audio frequency response table.
[0151] In some embodiments, the leaking audio data of each speaker and the target leakage audio response table are stored in the storage medium 420 of the wearable audio device 001.
[0152] In some embodiments, for the convenience of calculation, adjacent sub-bands can be fused to obtain sub-bands with a wider frequency range width. For example, when mapping the sound source channels, taking the left channel as an example, 64 sub-bands can be fused into 4 sub-bands, and the frequency ranges corresponding to the 4 sub-bands are 0 - 1KHz, 1KHz - 5KHz, 5KHz - 10KHz, and above 10KHz respectively. The wearable audio device 001 divides the target audio signal into 4 parts corresponding to the above 4 sub-bands through 3 crossover filters, that is, 4 sub-audio signals. When played through the speakers, the gains corresponding to the 4 sub-audio signals are determined based on the target leakage audio response table.
[0153] The gains corresponding to the 4 sub-audio signals are different. When the 4 sub-audio signals are switched and played, the gain switching and adjustment speed are slow. In this way, in the user's listening perception, the sound direction conversion caused by the gain transfer between different speakers is not easily perceived by the user, which is beneficial to maintaining the sound stability in the user's listening perception.
[0154] In some embodiments, the audio processing module 400 can process the target audio signal to be played within a time window to obtain the inherent volume of the target audio signal, and then determine the gain of the target audio signal based on the inherent volume of the target audio signal and the target leakage audio response table. If the inherent volume of the target audio signal is greater than the threshold determined by the target leakage audio response table, the audio processing module 400 will reduce the gain of the target audio signal. If the inherent volume of the target audio signal is less than the threshold determined by the target leakage audio response table, the audio processing module 400 will increase the gain of the target audio signal. For example, the audio processing module 400 can obtain the inherent volume of the target audio signal of a song to be played, and then determine the gain of the target audio signal based on the inherent volume of the target audio signal and the target leakage audio response table.
[0155] In some embodiments, technicians can design frequency response compensation algorithms for the sound emission characteristics of air-conduction speakers and bone-conduction speakers respectively to compensate the inherent frequency response curves of air-conduction speakers and bone-conduction speakers, so that the compensated frequency response curves are flat and consistent. Among them, the inherent frequency response curve can reflect the sound characteristics of the speaker hardware. The flatter the inherent frequency response curve, the higher the fidelity of the speaker when playing the target audio signal. When the inherent frequency response curve does not meet the requirements, the frequency response compensation algorithm can modify and compensate the inherent frequency response curve from software to make the compensated frequency response curve meet the requirements.
[0156] In some embodiments, technicians can perform a frequency response test on an air-conduction speaker unit to obtain the inherent frequency response curve of the air-conduction speaker. Specifically, technicians can perform a frequency response test on the air-conduction speaker unit using a microphone in an anechoic chamber environment to obtain the inherent frequency response curve of the air-conduction speaker unit.
[0157] In some embodiments, technicians can perform a frequency response test on the entire air-conduction speaker to obtain the inherent frequency response curve of the air-conduction speaker. Specifically, technicians can assemble the air-conduction speaker in the wearable audio device 001 as a whole, then fix the wearable audio device 001 through a simple bracket that can reduce the occlusion of the wearable audio device 001, and place the microphone at a position 1 cm away from the sound outlet hole of the air-conduction speaker. The air-conduction speaker can play white noise and stepped sweep tones. Technicians analyze the loudness and distortion of the playback signal after collecting the playback signal using the microphone, so as to obtain the inherent frequency response curve of the air-conduction speaker.
[0158] Among them, white noise refers to noise whose power spectral density is constant throughout the frequency domain, and all frequencies of white noise have the same energy density. The frequency of the stepped sweep tone is still constant within each time slice, but gradually increases between time slices, thus sweeping through the acoustic frequency range.
[0159] In some embodiments, technicians can perform a frequency response test on the air-conduction speaker under a simulated working scenario to obtain the inherent frequency response curve of the air-conduction speaker. Specifically, technicians can assemble the air-conduction speaker in the wearable audio device 001, then wear the wearable audio device 001 on an artificial head, and place the microphone in the ear of the artificial head. The air-conduction speaker can play white noise and stepped sweep tones. Technicians analyze the loudness and distortion of the playback signal after collecting the playback signal using the microphone, so as to obtain the inherent frequency response curve of the air-conduction speaker. The artificial head can simulate the structure of a user's head. This inherent frequency response curve can reflect the influence of the artificial head structure on sound propagation, that is, it can reflect the influence of the user's head structure on sound propagation, making the measured inherent frequency response curve closer to the user's listening perception.
[0160] In some embodiments, technicians can use a vibration tester to perform a frequency response test on a bone-conduction speaker unit to obtain the vibration condition of the bone-conduction speaker unit. Among them, the vibration condition of the bone-conduction speaker unit corresponds to the input data of the bone-conduction speaker. Based on the vibration condition of the bone-conduction speaker unit, technicians can determine the inherent frequency response curve of the bone-conduction speaker unit.
[0161] In some embodiments, a technician may assemble a bone conduction speaker in a wearable audio device 001. A tester (e.g., testers of different ages with rich listening experience) wears the wearable audio device 001 and sits at the central position of an anechoic chamber. A sound playback device such as a standard monitoring speaker or a tested open earphone is placed at a position 50 cm directly in front of the tester. The bone conduction speaker and the monitoring speaker sequentially play test audio signals such as single-frequency tone signals of each frequency, standard test voices, and standard test music. Among them, the audio signal of the bone conduction speaker is not processed, and the audio signal of the sound playback device allows the tester to manually adjust parameters such as gain and time delay through software, so that the listening sensations of the sound played by the bone conduction speaker and the sound played by the sound playback device are similar or consistent. Based on the audio signal of the sound playback device and the adjusted parameters, the technician can determine the inherent frequency response curve of the bone conduction speaker.
[0162] Furthermore, if there are multiple testers, the technician can determine multiple inherent frequency response curves corresponding to different testers, and a more accurate inherent frequency response curve of the bone conduction speaker can be obtained after weighted averaging the multiple inherent frequency response curves.
[0163] In addition, by combining the above two methods for obtaining the inherent frequency response curve of the bone conduction speaker, the technician can obtain the mapping relationship between the input data of the bone conduction speaker and the inherent frequency response curve. Then, based on this mapping relationship and the input data of the bone conduction speaker unit, the audio processing module 400 can determine the inherent frequency response curve of the bone conduction speaker unit.
[0164] Figure 4 The flowchart of adjusting the signal intensity ratio provided according to some embodiments of the present specification is shown. When the user is using the wearable audio device 001 and adjusts the volume from a high volume to a medium volume or a lower volume, it means that the user may be in a quiet environment or enter a quiet environment from a noisy environment. At this time, the focus of playback should be on controlling sound quality and sound leakage. Accordingly, as Figure 4 shown, S400 can be further described as:
[0165] S421: Receive a tuning instruction issued by the user, indicating that the total volume is adjusted to be lower than a first preset value;
[0166] The user can issue a tuning instruction through the volume key 501. The total volume being lower than the first preset value means that the total volume is at a medium volume or a lower volume, and at this time, the sound leakage of the air conduction speaker is low.
[0167] S422: Based on the source audio, adjust the allocation ratio of the first audio signal and the second audio signal, and reduce the proportion of the frequency signals with a high sound intensity ratio compared to the source audio in the first audio signal and the second audio signal before.
[0168] Specifically, the audio processing module 400 adjusts the proportions of the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal in the first audio signal based on the volume fluctuations of the sound source, and adjusts the proportions of the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal in the second audio signal.
[0169] This can make the proportions of the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal in the first audio signal close to or equal to those in the sound source audio, and make the proportions of the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal in the second audio signal close to or equal to those in the sound source audio. More specifically:
[0170] Adjusting the proportions of the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal in the first audio signal includes: adjusting the first low-frequency gain, the first mid-frequency gain, and the first high-frequency gain respectively applied to the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal.
[0171] Adjusting the proportions of the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal in the second audio signal includes: adjusting the second low-frequency gain, the second mid-frequency gain, and the second high-frequency gain respectively applied to the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal.
[0172] Figure 5 The flowchart of adjusting the signal strength ratio provided according to some embodiments of the present specification is shown. When the user is using the wearable audio device 001 and adjusts the volume from a small volume or a medium volume to a large volume, it indicates that the user may not be able to hear clearly. At this time, the focus of playback should be to increase the volume first, and control the sound quality and sound leakage second. Accordingly, as Figure 5 shown, S400 can be further described as:
[0173] S431: Receive the tuning instruction issued by the user, indicating to adjust the total volume to be higher than the second preset value;
[0174] The user can issue a tuning instruction through the volume key 501. The total volume being higher than the second preset value means that the total volume is at a large volume. At this time, the user has a relatively loose control over the sound leakage of the air conduction speaker.
[0175] S432: Based on the volume fluctuations of the sound source, adjust the allocation ratio of the first audio signal and the second audio signal, so that the ratio is adjusted in the direction of maximizing the gain of the bone conduction speaker module and the air conduction speaker module in each frequency band.
[0176] In some embodiments, the audio playback method P100 further includes: increasing the preset range of the sound leakage of the speaker module. At this time, the user has a relatively loose control over the sound leakage of the air conduction speaker, and a higher sound leakage volume is also allowed.
[0177] In some embodiments, the audio playback method P100 further includes: adjusting the allocation ratio of the first audio signal and the second audio signal based on the frequency response law of the air conduction speaker module 200 and the frequency response law of the bone conduction speaker module 300, so as to adjust the allocation ratio in the direction of maximizing the gain of the bone conduction speaker module 300 and the air conduction speaker module 200 in each frequency band.
[0178] In some embodiments, the first audio signal includes a first low-frequency signal, a first mid-frequency signal, and a first high-frequency signal; the second audio signal includes a second low-frequency signal, a second mid-frequency signal, and a second high-frequency signal;
[0179] Wherein adjusting the allocation ratio of the first audio signal and the second audio signal includes: adjusting the first low-frequency gain, the first mid-frequency gain, and the first high-frequency gain respectively applied to the first low-frequency signal, the first mid-frequency signal, and the first high-frequency signal; and adjusting the second low-frequency gain, the second mid-frequency gain, and the second high-frequency gain respectively applied to the second low-frequency signal, the second mid-frequency signal, and the second high-frequency signal.
[0180] Figure 6 The flowchart of adjusting the signal strength ratio provided according to some embodiments of the present specification is shown. In some embodiments, as Figure 6 shown, based on the volume fluctuation of the sound source, respectively controlling the proportions of the low-frequency signal, the mid-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal includes:
[0181] S433: Receive a tuning instruction issued by the user, indicating that the total volume is adjusted to be higher than the second preset value;
[0182] S434: Increase the preset range of sound leakage of the speaker module;
[0183] S435: Further adjust the distribution of different frequency amplitudes in the first audio signal, so that the proportion of the first mid-frequency signal in the first audio signal is further increased;
[0184] S436: Further adjust the distribution of different frequency amplitudes in the second audio signal described above, so that the proportion of the second mid-frequency signal is higher than the proportion of the second low-frequency signal and the proportion of the second high-frequency signal.
[0185] Specifically, the user can adjust the volume of the wearable audio device 001 through the volume button 501. The volume range that the volume button 501 can adjust may include a first volume range and a second volume range, and the volume of the first volume range is less than the volume of the second volume range. For example, the low volume and medium volume of the wearable audio device 001 are in the first volume range, and the high volume of the wearable audio device 001 is in the second volume range.
[0186] The audio processing module 400 can obtain the position of the volume adjusted by the volume key 501 and determine whether the volume is in the first volume range or the second volume range.
[0187] When the volume adjusted by the user through the volume key 501 is in the first volume range, that is, the total volume is lower than the first preset value, indicating that the user needs a smaller volume for the wearable audio device 001 to play sound, the frequency response compensation algorithm can be used to improve the sound quality fidelity. At this time, whether for the air conduction speaker or the bone conduction speaker, the specific steps of the frequency response compensation algorithm are as follows: The audio processing module 400 can lower the high loudness frequency points of the inherent frequency response curve, that is, control the maximum volume of each frequency point of the played content, so as to improve the overall flatness of the speaker frequency response curve. Then, the audio processing module 400 can apply a stable time-domain gain to the overall frequency response curve to adjust the volume. Among them, the time-domain gain can be determined by the audio processing module 400 based on the volume adjusted by the volume key 501 and is positively correlated with the volume adjusted by the volume key 501; the time-domain gain can also be determined by the audio processing module 400 based on the volume of the target audio signal and is positively correlated with the volume of the target audio signal.
[0188] When the volume adjusted by the user through the volume key 501 is in the second volume range, that is, the total volume is higher than the second preset value, or when the user adjusts the volume to a larger value through the volume key 501, indicating that the user needs a larger volume for the wearable audio device 001 to play sound, the frequency response compensation algorithm can be used to increase the volume. At this time, whether for the air conduction speaker or the bone conduction speaker, the specific steps of the frequency response compensation algorithm are as follows: On the premise that the hardware playback ability of the speaker allows the signal to be played without distortion, the audio processing module 400 can adjust the gain of each frequency point to maximize the gain of each frequency point. At this time, the low loudness frequency points are pulled up, and the sound quality fidelity of the played sound will be reduced. Among them, the fidelity refers to the degree of restoration of the sound played by the speaker to the sound of the target audio signal.
[0189] In some embodiments, when the frequency response compensation algorithm is used to increase the volume, the following steps are further included: The audio processing module 400 amplifies the gain of the sound frequency bands sensitive to the human ear, for example, amplifies the volume of the frequency band of 1KHz - 5KHz.
[0190] In some embodiments, the wearable audio device 001 includes a microphone module 600. When the frequency response compensation algorithm is used to increase the volume, the following steps are further included: The audio processing module 400 obtains the loudness of the ambient sound in real time through the microphone module 600 as the noise loudness, and calculates the signal-to-noise ratio (SNR) between the sound played by the speaker and the noise in real time. Based on the signal-to-noise ratio, the audio processing module 400 can determine whether the user can clearly hear the sound played by the speaker. If the audio processing module 400 determines that the user cannot clearly hear the sound played by the speaker, the volume will be increased so that the signal-to-noise ratio is not less than 10dB.
[0191] Further, the audio processing module 400 can calculate the signal-to-noise ratio of the sound played by the speaker at each frequency point or frequency band and the noise, so as to estimate the frequency band that the user cannot hear clearly, and enhance the volume of the corresponding frequency band so that the signal-to-noise ratio is not less than 10 dB. At this time, the signal may be distorted, but the user's perception is not obvious in a noisy environment.
[0192] In some embodiments, when the wearable audio device 001 plays sound in an application scenario, the audio processing module 400 can use a sound effect algorithm to improve the sound effect of each speaker.
[0193] In some embodiments, the sound effect algorithm may include at least one of a gain control algorithm, a frequency response control algorithm, harmonic enhancement, reverberation, etc. The audio processing module 400 can use these sound effect algorithms to improve the sound effect of the air conduction speaker.
[0194] In some embodiments, when the audio processing module 400 uses a sound effect algorithm to improve the sound effect of each speaker, it is necessary to balance between improving the sound effect of each speaker and reducing the sound leakage of each speaker. The audio processing module 400 can use a sound effect-sound leakage balance algorithm to balance between improving the sound effect of each speaker and reducing the sound leakage of each speaker. The following content gives the specific steps of the sound effect-sound leakage balance algorithm.
[0195] Figure 7 The flowchart of adjusting the signal strength ratio provided according to some embodiments of the present specification is shown. In some embodiments, as Figure 7 shown, based on the volume fluctuation of the sound source, respectively controlling the proportion of the low-frequency signal, the medium-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal includes:
[0196] S437: Receive a tuning instruction issued by the user, indicating that the total volume is adjusted to be higher than the second preset value;
[0197] S438: Obtain the ambient sound of the audio device through the microphone module on the audio device;
[0198] S439: Regard the ambient sound as noise and calculate the signal-to-noise ratio of the total volume in each frequency band;
[0199] S440: Adjust the volume corresponding to the target frequency band in the first audio signal and / or the second audio signal, where the target frequency band is the frequency band with a signal-to-noise ratio higher than the preset value.
[0200] In some embodiments, the audio playback method P100 further includes: increasing the volume of the target frequency band to a signal-to-noise ratio not less than 10 dB.
[0201] When the wearable audio device 001 plays sound in an application scenario, the audio processing module 400 can collect the environmental sound signal of the application scenario through the microphone of the microphone module 600, calculate the long-term steady-state noise level based on the environmental sound signal, and then determine the control level for sound leakage.
[0202] Specifically, the audio processing module 400 can calculate the ratio of the default volume when the target audio signal is played to the long-term steady-state noise level to obtain the estimated signal-to-noise ratio. If the value of the estimated signal-to-noise ratio is greater than the preset threshold, it means that the application scenario is a relatively quiet high-signal-to-noise ratio scenario, and the control of sound leakage is relatively strict. If the value of the estimated signal-to-noise ratio is less than the preset threshold, it means that the application scenario is a relatively noisy low-signal-to-noise ratio scenario, and the control of sound leakage is relatively loose.
[0203] The audio processing module 400 can perform sound leakage control and frequency response compensation on the target audio signal in different frequency bands, which is beneficial to improving the effect of sound leakage control and frequency response compensation on the target audio signal. The following introduces the steps for the audio processing module 400 to divide frequency bands.
[0204] The audio processing module 400 can divide the frequency of the target audio signal into multiple frequency bands, such as 3 - 5 frequency bands, and then calculate the signal-to-noise ratio level of each frequency band. Among them, the number of divided frequency bands depends on the computing power of the audio processing module 400. The more the number of divided frequency bands, the greater the computational amount of the audio processing module 400, but the more precise the control of the sound perception.
[0205] Adjacent frequency bands are divided by the crossover point. For example, if two adjacent frequency bands are 500Hz - 1kHz and 1kHz - 5kHz respectively, the crossover point is 1kHz. During the process of dividing frequency bands, the audio processing module 400 can dynamically adjust the frequency of the crossover point, that is, the interval of the frequency band is dynamically changing.
[0206] The basis for the audio processing module 400 to determine the final crossover point is: to minimize the variance of the signal-to-noise ratio of the frequency points within the frequency band.
[0207] For example, the audio processing module 400 can dynamically adjust the position of the crossover point a total of n times, that is, divide the frequency band a total of n times. For example, n = 10, 50 or 100. Each frequency band has multiple frequency points. Among the n times, the audio processing module 400 can select the crossover point corresponding to the time when the variance of the signal-to-noise ratio of the frequency points within the frequency band is the smallest as the final crossover point.
[0208] Among them, the variance of the signal-to-noise ratio of the frequency points within the frequency band refers to the variance of the signal-to-noise ratio of different frequency points within each frequency band. The smaller the variance of the signal-to-noise ratio of the frequency points, the flatter the frequency response curve in this frequency band, and the better the effect of sound leakage control and frequency response compensation for this frequency band.
[0209] When the final frequency division point is determined, in the frequency domain, the audio processing module 400 can divide the target audio signal into multiple sub-audio source signals according to different frequencies, and use the multiple sub-audio source signals as the input data of the air conduction speaker and the bone conduction speaker respectively. For example, in the time domain, the audio processing module 400 can divide the target audio signal into multiple sub-audio source signals through a crossover filter according to different frequencies.
[0210] Based on the long-term steady-state noise level, the audio processing module 400 can determine the respective gains of the multiple sub-audio source signals when played by the speakers. Specifically, the audio processing module 400 increases the gain of the sub-audio source signal whose frequency coincides with the long-term steady-state noise level, and decreases the gain of the sub-audio source signal whose frequency does not coincide with the long-term steady-state noise level. For example, if the long-term steady-state noise level is 3kHz - 5kHz, the audio processing module 400 increases the gain of the sub-audio source signal whose frequency coincides with 3kHz - 5kHz, and decreases the gain of the sub-audio source signal whose frequency does not coincide with 3kHz - 5kHz. In this way, on the one hand, the playback sound of the target audio signal will not be covered by environmental noise, which is beneficial to the sound of the target audio signal being clearly heard, and on the other hand, it is beneficial to reduce the risk of sound leakage when the target audio signal is played.
[0211] Furthermore, the greater the difference between the frequency of the sub-audio source signal and the long-term steady-state noise level, the greater the amplitude by which the audio processing module 400 decreases the gain.
[0212] In some embodiments, when the air conduction speaker plays the sub-audio source signal, the risk of sound leakage can be reduced through the following steps. The audio processing module 400 can compare each sub-audio source signal with a preset first frequency response upper limit curve, and adjust the gain according to the comparison result. Among them, the first frequency response upper limit curve can be used as a criterion for judging whether there is sound leakage when the target audio signal is played by the air conduction speaker.
[0213] For example, when the frequency response of the sub-audio source signal is less than or equal to the first frequency response upper limit curve, the gain of the air conduction speaker for the sub-audio source signal remains unchanged at the default value; when the frequency response of the sub-audio source signal is greater than the first frequency response upper limit curve, the gain of the air conduction speaker for the sub-audio source signal is decreased, and at the same time, the decreased part of the gain is applied to the input signal of the bone conduction speaker module, so that the total volume of the wearable audio device 001 remains unchanged.
[0214] In some embodiments, the audio playback method P100 further includes: based on the different distances of the bone conduction speaker module 300 and the air conduction speaker module 200 from the auditory nerve, applying a corresponding phase difference between the first audio signal and the second audio signal.
[0215] Since the sound propagation paths of the bone conduction speaker and the air conduction speaker are different. In order to make the sounds of the bone conduction speaker and the air conduction speaker heard by the user match, it is necessary to estimate the time difference between the sounds of the bone conduction speaker and the air conduction speaker on their respective propagation paths, and apply a corresponding phase difference between the first audio signal and the second audio signal to compensate for the time difference.
[0216] In some embodiments, after the audio processing module 400 compensates for the time difference using the phase difference, the user can hear the sounds of the bone conduction speaker and the air conduction speaker simultaneously.
[0217] Furthermore, based on information such as the age, gender, facial pictures, and daily wearing habits filled in by the user, the audio processing module 400 can correct the time difference. The audio processing module 400 can apply the corrected time difference to the input data of the bone conduction speaker through a time delay filter to compensate for the time difference.
[0218] In some embodiments, when the audio processing module 400 needs to adjust the sound effect of the bone conduction speaker, it can first determine the desired frequency response curve of the bone conduction speaker from the expected sound effect of the bone conduction speaker, then obtain the sound effect processing data from the desired frequency response curve of the bone conduction speaker through the above mapping relationship, and then input the sound effect processing data into the bone conduction speaker.
[0219] In some embodiments, when the bone conduction speaker plays a sub-audio source signal, the risk of sound leakage can be reduced through the following steps. The audio processing module 400 can compare each sub-audio source signal with a preset second upper frequency response curve, and adjust the gain according to the comparison result. Among them, the second upper frequency response curve can be used as a criterion for judging whether sound leakage occurs when the target audio signal of the bone conduction speaker is played.
[0220] For example, when the frequency response of the sub-audio source signal is less than or equal to the second upper frequency response curve, the gain of the bone conduction speaker for the sub-audio source signal remains the default value unchanged; when the frequency response of the sub-audio source signal is greater than the second upper frequency response curve, the gain of the bone conduction speaker for the sub-audio source signal is reduced, and at the same time, the reduced part of the gain is applied to the input signal of the bone conduction speaker module, so that the total volume of the wearable audio device 001 remains unchanged.
[0221] In some embodiments, the air conduction speaker can form a dipole sound source in front of and behind the diaphragm, or the air conduction speaker can drive two independent diaphragms to form a dipole sound source. The dipole sound source can form an inverse sound field structure and reduce sound leakage by emitting reverse sound waves.
[0222] In some embodiments, the air-conduction speaker module 200 includes at least one dipole speaker. Specifically, the air-conduction speaker can form a dipole sound source in front of and behind the diaphragm, or the air-conduction speaker can drive two independent diaphragms to form a dipole sound source. The dipole sound source can form an inverse sound field structure, and reduce sound leakage by emitting reverse sound waves. In this way, the dipole speaker can reduce the low-frequency sound leakage problem.
[0223] In some embodiments, the air-conduction speaker can have two diaphragms in the same direction. The two diaphragms in the same direction can share a voice coil, and sound can be emitted from the two diaphragms in the same direction, which helps to improve the volume and quality of the sound.
[0224] In some embodiments, the wearable audio device 001 can be provided with a sound leakage hole, and the reverse sound wave emitted from the sound leakage hole can interfere with and cancel the leaked sound wave.
[0225] In summary, the method and device provided in this application can use the bone-conduction speaker module 300 and the air-conduction speaker module 200 to play sound together, which is beneficial to improving the playback ability of the wearable audio device 001 in terms of volume, so that the overall volume of the wearable audio device 001 can still reasonably present the source audio. The working states of the bone-conduction speaker module 300 and the air-conduction speaker module 200 can be coordinated with each other, so as to make full use of the respective characteristics of the bone-conduction speaker module 300 and the air-conduction speaker module 200 to improve the sound effect of the wearable audio device 001 playing sound, and can also reduce the overall sound leakage of the wearable audio device 001 by reducing the sound leakage of the air-conduction speaker module 200.
[0226] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0227] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure can be presented only by way of example and is not restrictive. Although not explicitly stated here, those skilled in the art can understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0228] In addition, certain terms in this specification have been used to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. In addition, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0229] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, drawing, or its description. However, this does not mean that the combination of these features is necessary. It is entirely possible for those skilled in the art, when reading this specification, to mark out some of the devices as separate embodiments for understanding. That is to say, the embodiments in this specification can also be understood as the integration of multiple sub - embodiments. And it also holds when the content of each sub - embodiment contains fewer features than all the features of a single foregoing disclosed embodiment.
[0230] Every patent, patent application, published patent application, and other materials cited in this disclosure, such as articles, books, specifications, publications, documents, literature, etc. (excluding any historical examination documents related thereto), are hereby incorporated by reference for all purposes related to this disclosure, such as in the specification and claims of this disclosure. However, if there are any inconsistencies or conflicts between the descriptions, definitions, and / or terms used in the above - mentioned materials and those used in this disclosure, the descriptions, definitions, and / or terms used in this disclosure shall prevail.
[0231] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. An audio playback method for controlling sound leakage, characterized in that: Applied to wearable audio devices, including: Obtaining a target audio signal, wherein the target audio signal is a mapping of a source audio, and the source audio has source volume fluctuations; dividing the target audio signal into a first audio signal and a second audio signal; sending the first audio signal to a bone conduction speaker module in the audio device to emit bone conduction sound; sending the second audio signal to an air conduction speaker module in the audio device to emit air conduction sound; and Based on the sound source audio, the signal strengths of the first audio signal and the second audio signal are controlled respectively, so that the total volume of the air conduction sound and the bone conduction sound increases or decreases within a preset error following the fluctuation of the sound source volume, and at the same time, the sound leakage of the air conduction speaker module is controlled within a preset range.
2. The method according to claim 1, characterized in that in: The target audio signal includes a low-frequency signal, a medium-frequency signal and a high-frequency signal; The controlling the signal strengths of the first audio signal and the second audio signal respectively based on the sound source audio includes: controlling the proportions of the low-frequency signal, the intermediate-frequency signal and the high-frequency signal in the first audio signal and the second audio signal respectively based on the volume fluctuation of the sound source.
3. The method according to claim 2, characterized in that in: The first audio signal includes a first low-frequency signal, a first intermediate-frequency signal and a first high-frequency signal; the second audio signal includes a second low-frequency signal, a second intermediate-frequency signal and a second high-frequency signal; The controlling the proportions of the low-frequency signal, the intermediate-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal respectively based on the volume fluctuation of the sound source comprises: Based on the volume fluctuation of the sound source, a first gain and a second gain are respectively applied to the first audio signal and the second audio signal to adjust the distribution of different frequency amplitudes in the first audio signal and the second audio signal, so that the proportion of the first intermediate frequency signal is higher than the proportion of the first low frequency signal and the first high frequency signal, and the proportion of the second intermediate frequency signal is lower than the proportion of the second low frequency signal and the second high frequency signal.
4. The method according to claim 3, characterized in that The first gain and the second gain are determined based on a target leakage response table stored in a storage medium of the audio device.
5. The method according to claim 3, characterized in that in, The applying a first gain and a second gain to the first audio signal and the second audio signal respectively comprises: applying a first low-frequency gain, a first intermediate-frequency gain and a first high-frequency gain to the first low-frequency signal, the first intermediate-frequency signal and the first high-frequency signal respectively, and A second low-frequency gain, a second intermediate-frequency gain, and a second high-frequency gain are applied to the second low-frequency signal, the second intermediate-frequency signal, and the second high-frequency signal, respectively.
6. The method according to claim 5, characterized in that The intermediate frequency signal also includes an intermediate low frequency signal and an intermediate high frequency signal; Correspondingly, the first intermediate frequency signal includes a first intermediate low frequency signal and a first intermediate high frequency signal, and the first gain includes a first intermediate low frequency gain and a first intermediate high frequency gain; the second intermediate frequency signal includes a second intermediate low frequency signal and a second intermediate high frequency signal, and the second gain includes a second intermediate low frequency gain and a second intermediate high frequency gain; Correspondingly, applying the first intermediate frequency gain to the first intermediate frequency signal includes: applying the first intermediate low frequency gain to the first intermediate low frequency signal, and applying the first intermediate high frequency gain to the first intermediate high frequency signal; applying the second intermediate frequency gain to the second intermediate frequency signal includes: applying the second intermediate low frequency gain to the second intermediate low frequency signal, and applying the second intermediate high frequency gain to the second intermediate high frequency signal.
7. The method according to claim 6, characterized in that in, The applying a first gain and a second gain to the first audio signal and the second audio signal respectively based on the volume fluctuation of the sound source comprises: Based on the volume fluctuation of the sound source, the value of the first low-frequency gain, the value of the first intermediate-frequency gain, the value of the first high-frequency gain, the value of the second low-frequency gain, the value of the second intermediate-frequency gain and the value of the second high-frequency gain are controlled respectively to adjust the distribution of the first audio signal and the second audio signal at the low frequency, intermediate frequency and high frequency based on the volume fluctuation of the sound source.
8. The method according to claim 7, characterized in that in, The sound source audio includes low-frequency sound source audio, medium-frequency sound source audio and high-frequency sound source audio; the low-frequency sound source audio has a low-frequency sound source volume, the medium-frequency sound source audio has a medium-frequency sound source volume, and the high-frequency sound source audio has a high-frequency sound source volume; Correspondingly, the sound source volume fluctuations include low-frequency sound source volume fluctuations, medium-frequency sound source volume fluctuations, and high-frequency sound source volume fluctuations.
9. The method according to claim 8, characterized in that in, Based on the volume fluctuation of the sound source, applying a first gain and a second gain to the first audio signal and the second audio signal respectively comprises: When the volume of the intermediate frequency sound source increases, the second intermediate frequency gain is reduced and the first intermediate frequency gain is increased, so that a higher proportion of the energy of the intermediate frequency audio is allocated to the bone conduction speaker module.
10. The method according to claim 9, characterized in that in, Based on the fluctuation of the sound source volume, applying the first gain and the second gain to the first audio signal and the second audio signal respectively includes: increasing the second low-frequency gain and / or the second high-frequency gain, reducing the first low-frequency gain and / or the first high-frequency gain, so as to keep the total volume of the air conduction sound and the bone conduction sound increasing or decreasing within the preset error following the fluctuation of the sound source volume.
11. The method according to claim 8, characterized in that in, Based on the volume fluctuation of the sound source, applying a first gain and a second gain to the first audio signal and the second audio signal respectively comprises: When the volume of the high-frequency sound source exceeds a certain threshold, the second high-frequency gain is reduced and the first high-frequency gain is increased, so that a higher proportion of the high-frequency audio energy is allocated to the bone conduction speaker module.
12. The method according to claim 2 or 3, characterized in that: in, The controlling the proportions of the low-frequency signal, the intermediate-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal respectively based on the volume fluctuation of the sound source comprises: Receiving a tuning instruction from a user, instructing to adjust the total volume to be lower than a first preset value; Based on the sound source audio, the allocation ratio of the first audio signal and the second audio signal is adjusted to reduce the proportion of frequency signals with a higher sound intensity ratio than the sound source audio in the first audio signal and the second audio signal.
13. The method according to claim 12, characterized in that in, The audio source audio includes low-frequency audio source audio, intermediate-frequency audio source audio and high-frequency audio source audio; accordingly, the first audio signal includes a first low-frequency signal, a first intermediate-frequency signal and a first high-frequency signal; the second audio signal includes a second low-frequency signal, a second intermediate-frequency signal and a second high-frequency signal; The reducing the proportion of the frequency signals with a higher sound intensity ratio than the sound source audio in the first audio signal and the second audio signal comprises: Based on the volume fluctuation of the sound source, the proportions of the first low-frequency signal, the first intermediate-frequency signal and the first high-frequency signal in the first audio signal are adjusted, and the proportions of the second low-frequency signal, the second intermediate-frequency signal and the second high-frequency signal in the second audio signal are adjusted.
14. The method according to claim 13, characterized in that in, The adjusting the proportions of the first low-frequency signal, the first intermediate-frequency signal and the first high-frequency signal in the first audio signal comprises: adjusting a first low-frequency gain, a first intermediate-frequency gain and a first high-frequency gain respectively applied to the first low-frequency signal, the first intermediate-frequency signal and the first high-frequency signal; and Adjusting the proportions of the second low-frequency signal, the second intermediate-frequency signal and the second high-frequency signal in the second audio signal includes adjusting the second low-frequency gain, the second intermediate-frequency gain and the second high-frequency gain respectively applied to the second low-frequency signal, the second intermediate-frequency signal and the second high-frequency signal.
15. The method according to claim 2 or 3, characterized in that: Wherein, based on the volume fluctuation of the sound source, respectively controlling the proportions of the low-frequency signal, the intermediate-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal comprises: receiving a tuning instruction from a user, instructing to adjust the total volume to be higher than a second preset value; Based on the fluctuation of the sound source volume, the distribution ratio of the first audio signal and the second audio signal is adjusted so that the ratio is adjusted toward maximizing the gain of the bone conduction speaker module and the air conduction speaker module in each frequency band.
16. The method according to claim 15, characterized in that It also includes increasing the preset range of sound leakage of the speaker module.
17. The method according to claim 15, characterized in that Also includes: The allocation ratio of the first audio signal and the second audio signal is adjusted based on the frequency response law of the air conduction speaker module and the frequency response law of the bone conduction speaker module, so as to adjust the allocation ratio toward maximizing the gain of the bone conduction speaker module and the air conduction speaker module in each frequency band.
18. The method according to claim 17, characterized in that The first audio signal includes a first low-frequency signal, a first intermediate-frequency signal and a first high-frequency signal; the second audio signal includes a second low-frequency signal, a second intermediate-frequency signal and a second high-frequency signal; The adjusting the distribution ratio of the first audio signal and the second audio signal comprises: adjusting a first low-frequency gain, a first intermediate-frequency gain and a first high-frequency gain respectively applied to the first low-frequency signal, the first intermediate-frequency signal and the first high-frequency signal; as well as The second low-frequency gain, the second intermediate-frequency gain, and the second high-frequency gain respectively applied to the second low-frequency signal, the second intermediate-frequency signal, and the second high-frequency signal are adjusted.
19. The method according to claim 3, characterized in that Wherein, based on the volume fluctuation of the sound source, respectively controlling the proportions of the low-frequency signal, the intermediate-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal comprises: receiving a tuning instruction from a user, instructing to adjust the total volume to be higher than a second preset value; Improving the preset range of sound leakage of the speaker module; Further adjusting the distribution of different frequency amplitudes in the first audio signal so that the proportion of the first intermediate frequency signal in the first audio signal is further increased; The distribution of different frequency amplitudes in the second audio signal is further adjusted so that the proportion of the second intermediate frequency signal is higher than the proportion of the second low frequency signal and the second high frequency signal.
20. The method according to claim 2 or 3, characterized in that Wherein, based on the volume fluctuation of the sound source, respectively controlling the proportions of the low-frequency signal, the intermediate-frequency signal, and the high-frequency signal in the first audio signal and the second audio signal comprises: receiving a tuning instruction from a user, instructing to adjust the total volume to be higher than a second preset value; Acquiring the ambient sound of the audio device through a microphone module on the audio device; Treating the environmental sound as noise, and calculating the signal-to-noise ratio of the total volume in each frequency band; The volume corresponding to a target frequency band in the first audio signal and / or the second audio signal is adjusted, where the target frequency band is a frequency band where the signal-to-noise ratio is higher than a preset value.
21. The method of claim 20, wherein: Also includes: The volume of the target frequency band is increased until the signal-to-noise ratio is not less than 10 dB.
22. The method of claim 1, wherein: in, The air conduction speaker module includes at least one dipole speaker.
23. The method of claim 1, wherein: Also includes: Based on the different distances between the bone conduction speaker module and the air conduction speaker module and the auditory nerve, a corresponding phase difference is applied between the first audio signal and the second audio signal.
24. A wearable audio device comprising: A wearing bracket configured to be wearable on a user's head; An air conduction speaker module is disposed on the wearing bracket and faces the ear hole of the user and is at a preset distance from the ear hole of the user; A bone conduction speaker module is arranged on the wearing bracket, located behind the ear of the user and close to the skin of the user; An audio processing module is arranged on the wearing bracket and is communicatively connected with the air conduction speaker module and the bone conduction speaker module, and is configured to execute the audio playback method as claimed in any one of claims 1 to 23 during operation, and send the target audio signal to the air conduction speaker module and the bone conduction speaker module.
25. The device according to claim 24, characterized in that in, The wearing bracket is a glasses frame, including a left glasses leg and a right glasses leg; The air conduction speaker module includes at least one left air conduction speaker and at least one right air conduction speaker; as well as The bone conduction speaker module includes at least one left bone conduction speaker and at least one right bone conduction speaker; Wherein, the at least one left air conduction speaker and the at least one left bone conduction speaker are arranged on the left temple of the glasses; The at least one right air conduction speaker and the at least one right bone conduction speaker are arranged on the right glasses leg.
26. The device according to claim 24, characterized in that Also includes: A volume control module, which is disposed on the wearing bracket and is in communication connection with the audio processing module and is configured to receive a volume control signal input by a user; The audio processing module executes the audio playing method as claimed in any one of claims 14 to 21 based on the volume control signal.
27. The device according to claim 24, characterized in that Also includes: A volume control module, which is disposed on the wearing bracket and is in communication connection with the audio processing module and is configured to receive a volume control signal input by a user; as well as A microphone module, disposed on the wearing bracket and in communication with the audio processing module, configured to receive ambient sound around the device; The audio processing module further executes the audio playing method as claimed in any one of claims 14 to 23 based on the volume control signal.
28. The device according to any one of claims 24 to 26, characterized in that in, The audio processing module comprises: At least one storage medium storing at least one set of instructions for executing the audio playing method; At least one processor is communicatively connected with the at least one storage medium, the air conduction speaker module, the bone conduction speaker module and the volume control module, and executes the at least one set of instruction sets to execute the audio playback method during operation, thereby sending the target audio signal to the air conduction speaker module and the bone conduction speaker module.