Audio processing method, audio system, and audio device
By adaptively adjusting air conduction and bone conduction signals, and combining listening scene and environment recognition technology, the audio effect link of audio devices is optimized, solving the problem of poor listening experience in most scenarios of existing devices, and realizing a high-quality personalized audio experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-09-19
- Publication Date
- 2026-07-30
AI Technical Summary
Existing audio devices that combine bone conduction and air conduction technologies struggle to provide a high-quality listening experience in most scenarios, especially during audio processing, where the separation of high-frequency and low-frequency components results in poor sound quality.
By adaptively adjusting air conduction and bone conduction signals, the sound effect link is dynamically adjusted according to changes in the listening scenario and environment. Combined with listening scenario and environment recognition technology, the output of audio signals is optimized to adapt to the needs of different scenarios and environments.
It enables personalized, high-quality audio experiences in various scenarios and environments, enhances users' listening experience and motion performance, and improves the applicability and flexibility of audio devices.
Smart Images

Figure CN2025122672_30072026_PF_FP_ABST
Abstract
Description
Audio processing methods, audio systems and audio equipment
[0001] This application claims priority to Chinese Patent Application No. 202510126365.7, filed on January 27, 2025, entitled "Audio Processing Method, Audio System and Audio Device", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of electronic technology, and in particular to an audio processing method, an audio system, and an audio device. Background Technology
[0003] Currently, audio devices on the market that integrate bone conduction and air conduction technologies combine the advantages of bone conduction headphones (such as being less affected by external noise and having clearer audio in noisy environments) and air conduction headphones (such as having a better sense of presence and immersion), offering more comprehensive sound quality performance. They also take into account comfort, safety, and practicality, making them an ideal choice for modern consumers seeking a high-quality listening experience.
[0004] Currently, audio devices integrating bone conduction and air conduction technologies typically separate the audio signal to be played into high-frequency, mid-frequency, and low-frequency components during audio processing. The mid-high frequency components are played by vibrating speakers, while the mid-low frequency components are played by air-conducting speakers. However, this processing method means that the audio device can only provide a relatively good listening experience in a very few specific scenarios. In most scenarios, its listening effect is often unsatisfactory, failing to meet users' general demand for a high-quality audio experience. Summary of the Invention
[0005] In view of this, this application provides an audio processing method, an audio system, and an audio device. This audio processing method can provide a relatively good listening experience in most scenarios.
[0006] In a first aspect, this application provides an audio processing method, the method comprising: acquiring a first audio signal, the first audio signal being an audio signal in a multimedia file or a communication audio signal; acquiring a listening scenario; and outputting a second audio signal and a third audio signal; wherein the second audio signal is played through an air conduction audio playback unit, and the third audio signal is played through a bone conduction audio playback unit; when the first audio signal is the same, the second audio signal and the third audio signal output for different listening scenarios are different; and when the listening scenario is the same, the second audio signal and the third audio signal output for different first audio signals are different.
[0007] According to the statement that "when the first audio signal is the same, the second and third audio signals output for different listening scenarios are different", it can be seen that as the listening scenario changes, the output second and third audio signals also change accordingly. Therefore, this application can adaptively adjust the output air conduction signal (second audio signal) and bone conduction signal (third audio signal) according to the listening scenario. In this way, the audio device can be applied to a variety of scenarios and can provide a relatively good listening experience in most scenarios.
[0008] For example, "under the same listening scenario, the output second and third audio signals corresponding to different first audio signals are different" can be understood as the second and third audio signals both including the audio signals obtained by audio processing the first audio signal.
[0009] For example, the first audio signal may also be referred to as the downlink audio signal in some scenarios.
[0010] For example, multimedia files may include music files, video files (such as TV series files, movie files, etc.), streaming media files (which can refer to multimedia files transmitted in real time over a network, such as web page streaming media files, live streaming media files), conference audio and video streams, e-books, game files, screen recording files, animation files, call audio recording files, etc. In other words, the first audio signal can be the audio signal of music, the audio signal of dialogue or background sound from a TV series / movie, the audio signal of background sound or special effects from a game, the audio signal of dialogue or background sound from an animation, etc.
[0011] For example, the communication audio signal (the audio signal transmitted in a mobile communication system) can be a real-time audio signal obtained by acquiring voice and used for communication, such as a dial-up call audio signal. The communication audio signal can be a real-time audio signal obtained from other devices on the terminal device and acquired by those other devices.
[0012] For example, the listening scenario is the scenario in which the current audio device plays the audio signal; among which, the listening scenario includes, but is not limited to: sports scenario, movie scenario, music scenario, game scenario, call scenario, conference scenario, etc.
[0013] For example, an air-conduction audio playback unit can transmit sound waves through the air, which travel through the outer ear, tympanic membrane, and middle ear to reach the inner ear and are ultimately perceived by the auditory nerve; the air-conduction audio playback unit may include a speaker. A bone-conduction audio playback unit may include two propagation paths: one path transmits sound waves directly to the inner ear by vibrating the skull, bypassing the outer and middle ear; the other path transmits sound waves through the air, which travel through the outer ear, tympanic membrane, and middle ear to reach the inner ear; the bone-conduction audio playback unit may include a bone conduction vibrator.
[0014] According to the first aspect, the method further includes: acquiring the listening environment; wherein, when the first audio signal is the same and the listening scenario is the same, the second audio signal and the third audio signal output for different listening environments are different.
[0015] The statement that "under the same first audio signal and the same listening scenario, the output second and third audio signals differ depending on the listening environment" indicates that the output second and third audio signals change with the listening environment. In other words, this application can adaptively adjust the output air conduction signal (second audio signal) and bone conduction signal (third audio signal) according to the listening scenario, and also adaptively adjust the output air conduction signal (second audio signal) and bone conduction signal (third audio signal) according to the listening environment. Here, the listening scenario refers to the user's audio experience needs in a specific context; it focuses more on the user's behavior and purpose. The listening environment refers to the physical space where the user is located and its acoustic characteristics; it focuses more on the influence of the external environment on the audio experience. Therefore, the listening scenario and listening environment can jointly determine the final output audio signal; thus, by combining the listening scenario and listening environment, a more personalized and high-quality audio experience can be provided.
[0016] For example, the listening environment is the environment in which audio signals are played through audio devices. This listening environment is not limited to cinemas, elevators, shopping malls, subways, outdoor streets, mountains, living rooms, offices, etc.; it should be understood that the listening environment can be defined according to needs, and this application does not impose any restrictions on it.
[0017] According to the first aspect, or any implementation of the first aspect above, the third audio signal includes a bone conduction audio component and a bone conduction vibration component.
[0018] The bone conduction vibration component is used to control the intensity of the vibration of the bone conduction audio playback unit so that the user can feel the vibration. The bone conduction audio component is used to control the vibration of the bone conduction audio playback unit to generate sound so that the user can feel the sound. In other words, it can also provide cues to the user through vibration.
[0019] For example, in sports scenarios, when a smartwatch detects that a user's heart rate exceeds a pre-set heart rate threshold, or detects that a user's cadence is lower than a pre-set cadence threshold, it can generate bone conduction audio components and / or bone conduction vibration components to help improve the user's athletic performance.
[0020] According to the first aspect, or any implementation of the first aspect above, the second audio signal includes a first audio component and a third audio component, and the third audio signal includes a second audio component and a fourth audio component, wherein the second audio component and the fourth audio component are bone conduction audio components; the method further includes: acquiring a fourth audio signal, which is obtained by collecting ambient sound; wherein the first audio component and the second audio component are generated based on the first audio signal, and the third audio component and the fourth audio component are generated based on the fourth audio signal. This allows users to hear key ambient sounds or key human voices in the environment, enabling users to quickly locate the sound source in the real environment.
[0021] For example, key ambient sounds include key sounds in the current environment that may affect the user, such as car horns, traffic light signals, and alarms in an outdoor street environment. Key human voices include voiceprints of friends in the user's address book, user-defined keywords such as the user's name, and the voice of the person speaking to the user.
[0022] According to the first aspect, or any implementation of the first aspect above, the third audio signal further includes a sixth audio component, which is a bone conduction vibration component, generated based on the fourth audio signal. In this way, key ambient sounds in the ambient audio signal (i.e., the fourth audio signal) can be simulated with varying degrees of bone conduction vibration in the left and right earphones to indicate the location of the sound source to the user.
[0023] For example, when a user runs across an intersection and a car approaches from the left and honks its horn, the left bone conduction vibrator in the bone conduction audio playback unit can vibrate significantly stronger than the right bone conduction vibrator to indicate to the user that the horn is honking from the left.
[0024] According to the first aspect, or any implementation of the first aspect above, the third audio signal further includes an eighth audio component, which is a bone conduction vibration component, and the eighth audio component is generated based on the rhythm component in the first audio signal.
[0025] For example, in sports scenarios, drum beats in music that are at the same frequency as the user's stride can be converted into bone conduction vibration components; this can improve the sense of rhythm during exercise, enhance power and endurance, and improve the user's athletic performance.
[0026] According to the first aspect, or any implementation of the first aspect above, the method further includes: obtaining an audio effect link based on the listening scenario, or obtaining an audio effect link based on the listening scenario and the audio effect information input by the user; and performing audio effect processing on the first audio signal according to the audio effect link to obtain the second audio signal and the third audio signal.
[0027] The audio effects chain, also known as the audio effects algorithm chain, refers to the process of connecting multiple audio effects algorithms in a specific order during audio processing to gradually change or enhance the audio signal. It's important to note that configuring the audio effects chain involves configuring not only the required multiple audio effects algorithms and their connection order, but also the parameters of any individual audio effects algorithm. These algorithms include, but are not limited to: active noise reduction algorithms, dynamic range compression (DRC) algorithms, adaptive EQ (adaptive EQ) algorithms, adaptive dynamic frequency division algorithms, vibration enhancement algorithms, rhythm enhancement algorithms, key sound separation algorithms, vibration suppression algorithms, and sound leakage suppression algorithms.
[0028] In other words, this application adaptively switches the matching audio effect link according to the listening scenario (or according to the listening scenario and the audio effect information input by the user); by using different audio effect links to process the first audio signal, the proportion and listening experience of each component can be adjusted to ensure the best sound quality and the best experience; thereby ensuring the listening experience in multiple listening scenarios.
[0029] According to the first aspect, or any implementation of the first aspect above, the method further includes: obtaining an audio effect link based on the listening scenario and the listening environment, or obtaining an audio effect link based on the listening scenario, the listening environment and the audio effect information input by the user; and performing audio effect processing on the first audio signal according to the audio effect link to obtain the second audio signal and the third audio signal.
[0030] In other words, this application adaptively switches the matching sound effect link according to the listening scenario and listening environment (or according to the listening scenario, listening environment and user-input sound effect information); thereby ensuring the listening experience of multiple listening scenarios and listening environments.
[0031] It should be noted that this application does not protect a single audio algorithm or sound effect chain, but rather protects the ability to switch between different sound effect chains in different listening scenarios and environments to adjust the proportions of air conduction audio, bone conduction audio components, and bone conduction vibration components, thereby making the audio device suitable for different listening scenarios and environments, achieving different functions, and satisfying diverse listening experiences.
[0032] According to the first aspect, or any implementation of the first aspect above, the method further includes: performing sound effect processing on the first audio signal based on the audio sound effect link included in the sound effect link to obtain the first audio component and the third audio component; performing sound effect processing on the fourth audio signal based on the ambient sound effect link included in the sound effect link to obtain the second audio component and the fourth audio component; wherein the sound effect link is determined based on the listening scene and the listening environment, or the audio link is determined based on the listening scene, the listening environment and the sound effect information input by the user.
[0033] According to the first aspect, or any implementation of the first aspect above, the acquisition of the listening scenario includes: determining the listening scenario based on at least one of the user input scenario, the type of the first audio signal, or the device usage status.
[0034] In one possible implementation, the listening scenario can be determined based on the type of the first audio signal. For example, if the first audio signal is a video audio type, the listening scenario can be determined to be a movie scenario. As another example, if the first audio signal is a game audio type, the listening scenario can be determined to be a game scenario. Yet another example, if the first audio signal is a call audio type, the listening scenario can be determined to be a call scenario. And yet another example, if the first audio signal is a conference audio type, the listening scenario can be determined to be a conference scenario.
[0035] In one possible implementation, the listening scenario can be determined based on the device usage state. This device usage state can include the usage state of the terminal device (including the terminal device used to execute the audio processing method of this application and / or other terminal devices) and / or the usage state of the audio device. For example, if music is playing on the terminal device and the activity tracking function on a smartwatch is enabled, the listening scenario can be determined to be an activity scenario. As another example, if an incoming call or a pre-arranged call is answered on the terminal device, the listening scenario can be determined to be a call scenario. Yet another example is if an AR / VR headset is playing a TV series / movie in a video application, the listening scenario can be determined to be a movie scenario. Furthermore, if the extension module is a simultaneous interpretation module and the simultaneous interpretation module is enabled, the listening scenario can be determined to be a cross-language conversation scenario (such as a cross-language conference scenario, a cross-border shopping scenario, etc.).
[0036] In one possible implementation, the listening scenario can be determined based on the user's input scenario. Specifically, the user's input scenario can be used to determine the listening scenario.
[0037] According to the first aspect, or any implementation of the first aspect above, the acquisition of the listening environment includes: determining the listening environment based on at least one of a fourth audio signal, other sensor data, device usage status, or user input; wherein the fourth audio signal is obtained by collecting ambient sound.
[0038] In one possible implementation, the listening environment can be determined based on a fourth audio signal. This fourth audio signal is obtained by collecting ambient sound data and can also be referred to as an ambient sound audio signal. The fourth audio signal can be collected by a terminal device (including a terminal device executing the audio processing method of this application and / or other terminal devices (such as smartwatches)) or by an audio device. Exemplarily, the terminal device can perform environmental recognition based on the fourth audio signal to determine the listening environment. For example, the terminal device can extract key features from the fourth audio signal, such as the spectrum, Mel-frequency cepstral coefficients (MFCC), sound intensity, and frequency distribution. Subsequently, the extracted key features can be input into a classification model (which can be implemented by a pre-trained neural network), and the classification model outputs the listening environment.
[0039] In one possible implementation, the listening environment can be determined based on data from other sensors. This other sensor data can refer to data collected by sensors other than the microphone (such as Global Positioning System (GPS), accelerometer, gyroscope, barometer, light sensor, temperature sensor, magnetic field sensor, etc.). This other sensor data can be collected by the terminal device executing the audio processing method of this application and / or other terminal devices and / or audio devices. For example, the user's geographical location (e.g., living room, office, park) can be determined based on GPS data; subsequently, the listening environment can be determined based on the user's geographical location. Another example is determining the user's motion state (e.g., stationary, walking, running) based on data from an accelerometer; then, the listening environment (e.g., car, subway) can be determined based on the user's motion state. Yet another example is determining the rotation and orientation of the terminal device based on data from a gyroscope; then, the listening environment (e.g., elevator, stairs) can be determined based on the rotation and orientation of the terminal device. Finally, determining whether the user is in an indoor or outdoor environment can be determined based on data from a light sensor. For example, based on data collected by a magnetic field sensor, it can identify whether a user is near a specific device (such as a stereo or television); and further identify whether the user is in the living room or office.
[0040] One possible implementation is to determine the listening environment based on the device's usage status. For example, the listening environment can be determined based on the device's volume; high volume may indicate the user is in a noisy environment (such as an outdoor street or shopping mall), while low volume may indicate the user is in a quiet environment (such as a library or conference room). Alternatively, the listening environment can be determined based on the Wi-Fi information the device is connected to; for example, if the device is connected to office Wi-Fi, the listening environment is determined to be an office; if the device is connected to public Wi-Fi, the listening environment is determined to be a public place such as a coffee shop or airport. Furthermore, the listening scenario can be determined based on the application software being used on the device; for example, if video conferencing software is being used, the listening environment is determined to be an office; if navigation software is being used, the listening environment is determined to be outdoors.
[0041] In one possible implementation, the listening environment can be determined based on the user's input environment. Specifically, the user's input environment can be used to determine the listening environment.
[0042] According to the first aspect, or any implementation of the first aspect above, the second audio signal further includes a fifth audio component, which is generated based on the rhythmic component of the first audio signal. This enhances the rhythm in the first audio signal.
[0043] For example, the rhythmic components in the first audio signal may include, but are not limited to: drum beats, melodies, harmonies, bass, percussion instruments, rhythm guitars, keyboard instruments, string instruments, etc.
[0044] Secondly, this application provides an audio system, which includes an audio device and a terminal device, wherein the audio device includes an air conduction audio playback unit and a bone conduction audio playback unit;
[0045] The terminal device is used to acquire a first audio signal, which is an audio signal in a multimedia file or a communication audio signal; acquire a listening scenario; output a second audio signal to the air conduction audio playback unit; and output a third audio signal to the bone conduction audio playback unit.
[0046] Wherein, when the first audio signal is the same, the second and third audio signals output for different listening scenarios are different; when the listening scenario is the same, the second and third audio signals output for different first audio signals are different.
[0047] According to the second aspect, the audio device includes headphones and an expansion module, which includes at least one of an audio module, a health monitoring module, or an environmental monitoring module;
[0048] The headphones include the air conduction audio playback unit, and the expansion module includes the bone conduction audio playback unit.
[0049] It should be understood that an audio system implementing the second aspect or any of the above second aspects can execute the methods in the first aspect or any of the above first aspects, which will not be elaborated here.
[0050] Thirdly, this application provides an audio device, characterized in that the audio device includes an air conduction audio playback unit and a bone conduction audio playback unit.
[0051] The audio device is used to: acquire a first audio signal, which is an audio signal in a multimedia file or a communication audio signal; acquire a listening scenario; output a second audio signal to the air conduction audio playback unit; and output a third audio signal to the bone conduction audio playback unit.
[0052] Wherein, when the first audio signal is the same, the second and third audio signals output for different listening scenarios are different; when the listening scenario is the same, the second and third audio signals output for different first audio signals are different.
[0053] According to the third aspect, or any implementation of the third aspect above, the audio device includes headphones and an expansion module, the expansion module including at least one of an audio module, a health monitoring module, or an environmental monitoring module; the headphones include the air conduction audio playback unit, and the expansion module includes the bone conduction audio playback unit.
[0054] It should be understood that an audio device implementing any of the third or above third aspects can execute the methods in any of the first or above first aspects, which will not be elaborated here.
[0055] Fourthly, this application provides an electronic device, including: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, cause the electronic device to perform the method of the first aspect or any possible implementation thereof.
[0056] Fifthly, this application provides a chip including one or more interface circuits and one or more processors; the one or more processors receive or transmit data through the one or more interface circuits, and when the one or more processors execute computer instructions, cause the electronic device to perform the method in the first aspect or any possible implementation of the first aspect.
[0057] In a sixth aspect, this application provides a computer-readable storage medium storing a computer program that, when run on a computer or processor, causes the computer or processor to perform the method of the first aspect or any possible implementation thereof.
[0058] In a seventh aspect, this application provides a computer program product including computer instructions that, when executed by a computer or processor, cause the computer or processor to perform the method in the first aspect or any possible implementation thereof.
[0059] In this embodiment, the electronic device, computer-readable storage medium, computer program product, chip or audio device, audio system, etc. are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods provided above. Attached Figure Description
[0060] Figure 1A is a schematic diagram of an application scenario according to an embodiment of this application;
[0061] Figure 1B is a schematic diagram of another application scenario of the present application embodiment;
[0062] Figure 1C is a schematic diagram of another application scenario of the present application;
[0063] Figure 1D is a schematic diagram of another application scenario of the present application embodiment;
[0064] Figure 2A is a schematic diagram of an audio device according to an embodiment of this application;
[0065] Figure 2B is a schematic diagram of another audio device according to an embodiment of this application;
[0066] Figure 3A is a schematic diagram of an audio system 310 according to an embodiment of this application;
[0067] Figure 3B is a schematic diagram of another audio system 320 according to an embodiment of this application;
[0068] Figure 3C is a schematic diagram of another audio system 330 according to an embodiment of this application;
[0069] Figure 4 is a schematic diagram of an audio processing procedure 400 according to an embodiment of this application;
[0070] Figure 5A is a schematic diagram of a mobile phone interface according to an embodiment of this application;
[0071] Figure 5B is a schematic diagram of another mobile phone interface according to an embodiment of this application;
[0072] Figure 6 is a schematic diagram of another audio processing procedure 600 according to an embodiment of this application;
[0073] Figure 7A is a schematic diagram of another audio processing process 700 according to an embodiment of this application;
[0074] Figure 7B is a schematic diagram of an audio effect link according to an embodiment of this application;
[0075] Figure 8A is a schematic diagram of another audio processing process 800 according to an embodiment of this application;
[0076] Figure 8B is a schematic diagram of another audio effect link according to an embodiment of this application;
[0077] Figure 9A is a schematic diagram of another audio processing procedure 900 according to an embodiment of this application;
[0078] Figure 9B is a schematic diagram of another mobile phone interface according to an embodiment of this application;
[0079] Figure 9C is a schematic diagram of another mobile phone interface according to an embodiment of this application;
[0080] Figure 9D is a schematic diagram of another audio effect link according to an embodiment of this application;
[0081] Figure 10A is a schematic diagram of another mobile phone interface according to an embodiment of this application;
[0082] Figure 10B is a schematic diagram of another mobile phone interface according to an embodiment of this application;
[0083] Figure 10C is a schematic diagram of another mobile phone interface according to an embodiment of this application;
[0084] Figure 10D is a schematic diagram of another mobile phone interface according to an embodiment of this application;
[0085] Figure 10E is a schematic diagram of another mobile phone interface according to an embodiment of this application;
[0086] Figure 11 is a schematic diagram of the structure of a device provided in an embodiment of this application. Detailed Implementation
[0087] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0088] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0089] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.
[0090] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0091] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.
[0092] In the embodiments of this application, the modules / components shown in the framework diagram (or structural diagram or system diagram) are merely examples of this application. The actual framework (or structure or system) may include more or fewer modules / components than those shown in the diagram, or may have different component configurations. Furthermore, the various components / modules shown in the diagrams may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0093] Figure 1A is a schematic diagram of an application scenario according to an embodiment of this application. Figure 1A shows an audio device used in a call scenario.
[0094] Referring to Figure 1A, user A can dial user B's phone B using mobile phone A (user A's phone A and the audio device worn by user A have established a communication connection). After user B answers the call from mobile phone A, user A and user B can have a voice call. During the voice call, the audio device worn by user A can capture user A's voice to obtain audio signal A, and send audio signal A to mobile phone A, which then sends audio signal A to mobile phone B. After receiving audio signal A, mobile phone B, since its hands-free function is enabled, can play audio signal A through its internal speaker. Correspondingly, mobile phone B can capture user B's voice to obtain audio signal B, and send audio signal B to mobile phone A. After receiving audio signal B, mobile phone A can send audio signal B to the audio device, which can then play audio signal B. Using an audio device during a voice call has advantages such as freeing up hands, effectively reducing ambient noise, making calls clearer, and protecting privacy.
[0095] Figure 1B is a schematic diagram of another application scenario of the present application embodiment. Figure 1B shows the application of audio equipment in a sports scene.
[0096] Referring to Figure 1B, users can wear an audio device before or during running and establish a communication connection between the audio device and their mobile phone. Afterwards, users can play music, audiobooks, news, and other audio files on their mobile phones. The phone responds to the user's actions by sending the audio signals of the music, audiobooks, news, etc., to the audio device; the audio device can then play these audio signals. Using an audio device during exercise offers advantages such as freeing up hands, reducing external interference, and improving audio quality.
[0097] Figure 1C is a schematic diagram of another application scenario of this application embodiment. Figure 1C shows an audio device applied in a game scenario.
[0098] Referring to Figure 1C, users can wear an audio device before or during gameplay and establish a communication connection between the audio device and their mobile phone. During gameplay, the mobile phone can send audio signals of game sound effects (such as the opening theme song, battle background, button clicks, etc.) and teammates' voices to the audio device; the audio device can then play these audio signals. Using an audio device during gameplay offers advantages such as ensuring an immersive gaming experience, reducing external interference, and protecting privacy.
[0099] Figure 1D is a schematic diagram of another application scenario of this application embodiment. Figure 1D shows an audio device applied in a movie scene.
[0100] Referring to Figure 1D, users can wear an audio device before or during a movie and establish a communication connection between the audio device and their mobile phone. Then, during movie playback, the mobile phone can send the movie's audio signals (such as dialogue and background music) to the audio device; the audio device can then play the movie's audio signals. Using an audio device while watching a movie offers advantages such as ensuring an immersive viewing experience and providing clearer, higher-quality audio.
[0101] It should be understood that audio devices can also be used in other listening scenarios, such as music scenarios, meeting scenarios, audio production scenarios, etc.; this application does not limit this.
[0102] It should be understood that the mobile phones in the application scenarios of Figures 1A to 1D can be replaced with other devices, such as tablets, wearable devices (such as smartwatches, smart glasses, augmented reality (AR) / virtual reality (VR) devices, etc.), and this application does not impose any restrictions on them.
[0103] In one possible implementation, the audio device of this application can be a combined bone and air conduction audio device, that is, an audio device that includes an air conduction audio playback unit and a bone conduction audio playback unit.
[0104] Figure 2A is a schematic diagram of an audio device according to an embodiment of this application. The audio device in Figure 2A(1) is a head-mounted display device 21, the audio device in Figure 2A(2) is an ear-hook headphone 22, and the audio device in Figure 2A(3) is an in-ear headphone 23.
[0105] In one possible implementation, the audio device of this application may include headphones and a wearable audio device. The headphones include an air conduction audio playback unit, and the wearable audio device includes a bone conduction audio playback unit; or, the headphones include a bone conduction audio playback unit, and the wearable audio device includes an air conduction audio playback unit.
[0106] Figure 2B is a schematic diagram of another audio device according to an embodiment of this application. The audio device in Figure 2B(1) includes a head-mounted display device 24 and headphones 25, and the audio device in Figure 2B(2) includes smart glasses 26 and headphones 27.
[0107] In one possible implementation, the audio device of this application (as shown in Figures 1A to 1D) may include: headphones 29 and an expansion module 28, as shown in Figure 2B(3). The headphones are detachably connected to the expansion module via a connecting component. The headphones include an air conduction audio playback unit, and the expansion module includes a bone conduction audio playback unit; or, the headphones include a bone conduction audio playback unit, and the expansion module includes an air conduction audio playback unit.
[0108] For example, the headphones can be air-conducting headphones (or air-conducting earphones). These air-conducting headphones can be wired or wireless. For example, the form of the air-conducting headphones can be, for example, in-ear, earbud, over-ear, ear-hook, open-back, true wireless, etc., and this application does not impose any limitations on this.
[0109] For example, the extension module may include a module capable of playing audio independently and capable of playing audio in conjunction with headphones, and / or a module that can assist headphones in playing audio.
[0110] For example, the extension module may include at least one of an audio module, an environmental monitoring module, or a health monitoring module. It should be understood that the extension module may also be other types of modules, and this application does not limit this.
[0111] For example, an audio module can refer to a component that has the function of processing and playing audio signals as well as the function of acquiring audio signals, such as bone conduction headphones, hearing aids, artificial intelligence modules (such as voice assistant modules, translators or simultaneous interpreters (simultaneous interpreting modules)), etc.
[0112] For example, a health monitoring module can be used to monitor human health status and may include sensors for collecting health data, such as heart rate sensors, body temperature sensors, accelerometers, gyroscopes, etc.
[0113] For example, the environmental monitoring module can be used to monitor the environmental status and may include sensors for collecting environmental data, such as infrared detectors, gas detectors, radiation detectors, audio acquisition modules, image acquisition modules, etc.
[0114] It should be understood that the audio module, environmental monitoring module, or health monitoring module are all detachable modules.
[0115] For example, the earphone 29 and the expansion module 28 in Figure 2B(3) are connected by an ear hook bracket; it should be understood that the earphone 29 and the expansion module 28 can also be connected by other types of brackets such as earring brackets, neck brackets, and headband brackets; in addition, the earphone 29 and the expansion module 28 can also be connected by magnetic force, which is not limited in this application.
[0116] In one possible implementation, the audio device may include headphones and playback devices such as speakers, MP4 players, and televisions.
[0117] To address the problem that existing audio devices can only provide a relatively good listening experience in a very few specific scenarios, this application provides an audio system (including an audio device) and an audio processing method, so that the audio device can provide a relatively good listening experience in most scenarios.
[0118] Figure 3A is a schematic diagram of an audio system 310 according to an embodiment of this application.
[0119] The audio system 310 in Figure 3A may include a terminal device 41 and an audio device 42. Exemplarily, the terminal device 41 may include a scene recognition module 411, a sound effect link configuration module 412, an environment recognition module 413, a sound effect processing module 414, and a first communication module 415. The environment recognition module 413 is optional. Exemplarily, the audio device 42 may include a playback monitoring module 421, a playback module 422, and a second communication module 423; the playback module 422 may include an air conduction audio playback unit 4221 and a bone conduction audio playback unit 4222. The playback monitoring module 421 is optional.
[0120] For example, the scene recognition module 411 is used to recognize the listening scene.
[0121] For example, the audio effect link configuration module 412 is used to configure the audio effect link. The audio effect link can also be called an audio effect chain, audio effect algorithm chain, etc. An audio effect link refers to the processing flow in which multiple audio effect algorithms are connected in a specific order during audio processing to gradually change or enhance the audio signal. It should be noted that, in configuring the audio effect link, in addition to configuring the required multiple audio effect algorithms and their connection order, the parameters of any one of these audio effect algorithms can also be configured. These audio effect algorithms include, but are not limited to: active noise reduction algorithms, dynamic range compression (DRC) algorithms, adaptive EQ (Adaptive Equalization) algorithms, adaptive dynamic frequency division algorithms, vibration enhancement algorithms, rhythm enhancement algorithms, key sound separation algorithms, vibration suppression algorithms, and sound leakage suppression algorithms.
[0122] For example, the environment recognition module 413 is used to recognize the listening environment.
[0123] For example, the sound effect processing module 414 is used to call the sound effect algorithm to perform sound effect processing on the audio signal.
[0124] For example, the sound playback module 422 is used to play audio signals. The air conduction audio playback unit 4221 can transmit sound waves through the air, which travel through the outer ear, tympanic membrane, and middle ear to reach the inner ear and are ultimately perceived by the auditory nerve; for example, the air conduction audio playback unit 4221 may include a speaker. The bone conduction audio playback unit 4222 may include two propagation paths. One propagation path transmits sound waves directly to the inner ear by vibrating the skull, bypassing the outer and middle ear; the other propagation path transmits sound waves through the air, which travel through the outer ear, tympanic membrane, and middle ear to reach the inner ear; for example, the bone conduction audio playback unit may include a bone conduction vibrator.
[0125] For example, the playback monitoring module 421 is used to collect the audio signal played by the playback module 422 and feed the collected audio signal back to the sound effect processing module 414 and the sound effect link configuration module 412, so that the sound effect processing module 414 can optimize the sound effect algorithm and the sound effect link configuration module 412 can adjust the sound effect link.
[0126] For example, the first communication module 415 and the second communication module 423 are used to send and receive data.
[0127] For example, the audio device 42 may also include sensors. The audio device 42 may feed back the sensing data collected by the sensors to the scene recognition module 411 to provide a basis for scene recognition by the scene recognition module 411; and may also feed back the sensing data to the environment recognition module 413 to provide a basis for environment recognition by the environment recognition module 413.
[0128] Figure 3B is a schematic diagram of another audio system 320 according to an embodiment of this application.
[0129] The audio system 320 in Figure 3B may include a terminal device 43 and an audio device 44. Exemplarily, the terminal device 43 may include a scene recognition module 431, a sound effect link configuration module 432, an environment recognition module 433, and a first communication module 434. The environment recognition module 433 is optional. Exemplarily, the audio device 44 may include a playback monitoring module 441, a second communication module 442, a sound effect processing module 443, and a playback module 444; the playback module 444 may include an air conduction audio playback unit 4441 and a bone conduction audio playback unit 4442. The playback monitoring module 441 is optional.
[0130] The functions of each module / unit in Figure 3B are similar to those of each module / unit in Figure 3A, and will not be repeated here.
[0131] Figure 3C is a schematic diagram of another audio system 330 according to an embodiment of this application.
[0132] The audio system 330 in Figure 3C may include an audio device 45. Exemplarily, the audio device 45 may include a scene recognition module 451, a sound effect link configuration module 452, an environment recognition module 453, a sound effect processing module 454, a playback module 455, and a playback monitoring module 456; the playback module 455 may include an air conduction audio playback unit 4551 and a bone conduction audio playback unit 4552. The playback monitoring module 456 and the environment recognition module 453 are optional modules.
[0133] The functions of each module / unit in Figure 3C are similar to those of each module / unit in Figure 3A, and will not be repeated here.
[0134] For example, the scene recognition module, the sound effect link configuration module, the environment recognition module, the sound effect processing module, and the playback module can jointly execute the audio processing method of this application.
[0135] It should be understood that the scene recognition module, sound effect link configuration module, environment recognition module, and sound effect processing module can be deployed on the terminal device, on the audio device, or distributed on the terminal device and the audio device. This application does not impose any restrictions on this.
[0136] It should be understood that the terminal device can refer to a device other than the audio device mentioned above.
[0137] It should be understood that the terminal devices and audio devices in Figures 3A to 3C may also include audio acquisition modules (such as microphones or vibration pick-up (VPU)) for acquiring audio signals.
[0138] The audio processing method of this application will be explained below using Figure 3A as an example.
[0139] Figure 4 is a schematic diagram of an audio processing procedure 400 according to an embodiment of this application.
[0140] S401, acquire the first audio signal, which is an audio signal from a multimedia file or a communication audio signal.
[0141] For example, the first audio signal may also be referred to as the downlink audio signal in some scenarios.
[0142] For example, multimedia files may include music files, video files (such as TV series files, movie files, etc.), streaming media files (which can refer to multimedia files transmitted in real time over a network, such as web page streaming media files, live streaming media files), conference audio and video streams, e-books, game files, screen recording files, animation files, call audio recording files, etc. In other words, the first audio signal can be the audio signal of music, the audio signal of dialogue or background sound from a TV series / movie, the audio signal of background sound or special effects from a game, the audio signal of dialogue or background sound from an animation, etc.
[0143] For example, the terminal device may read the first audio signal from a locally pre-stored multimedia file, or it may obtain a multimedia file from another device to read the first audio signal from it.
[0144] For example, the communication audio signal (the audio signal transmitted in a mobile communication system) can be a real-time audio signal obtained by acquiring voice and used for communication, such as a dial-up call audio signal. The communication audio signal can be a real-time audio signal obtained from other devices on the terminal device and acquired by those other devices.
[0145] S402, Obtain the listening scene.
[0146] For example, S402 can be performed by the scene recognition module 411 in FIG3A.
[0147] For example, the listening scenario is the scenario in which the current audio device plays the audio signal; among which, the listening scenario includes, but is not limited to: sports scenario, movie scenario, music scenario, game scenario, call scenario, conference scenario, etc.
[0148] In one possible implementation, the listening scenario can be determined based on the type of the first audio signal. For example, if the first audio signal is a video audio type, the listening scenario can be determined to be a movie scenario. As another example, if the first audio signal is a game audio type, the listening scenario can be determined to be a game scenario. Yet another example, if the first audio signal is a call audio type, the listening scenario can be determined to be a call scenario. And yet another example, if the first audio signal is a conference audio type, the listening scenario can be determined to be a conference scenario.
[0149] In one possible implementation, the listening scenario can be determined based on the device usage state. This device usage state can include the usage state of the terminal device (including the terminal device used to execute the audio processing method of this application and / or other terminal devices) and / or the usage state of the audio device. For example, if music is playing on the terminal device and the activity tracking function on a smartwatch is enabled, the listening scenario can be determined to be an activity scenario. As another example, if an incoming call or a pre-arranged call is answered on the terminal device, the listening scenario can be determined to be a call scenario. Yet another example is if an AR / VR headset is playing a TV series / movie in a video application, the listening scenario can be determined to be a movie scenario. Furthermore, if the extension module is a simultaneous interpretation module and the simultaneous interpretation module is enabled, the listening scenario can be determined to be a cross-language conversation scenario (such as a cross-language conference scenario, a cross-border shopping scenario, etc.).
[0150] In one possible implementation, the listening scenario can be determined based on the user's input scenario. Specifically, the user's input scenario can be used to determine the listening scenario.
[0151] Figure 5A is a schematic diagram of a mobile phone interface according to an embodiment of this application.
[0152] Referring to Figure 5A(1), exemplarily, 501 is the main interface of the mobile phone, which includes one or more controls, including but not limited to: application icons (e.g., application icon of browser application, application icon of settings application 502), etc.
[0153] Referring again to Figure 5A(1), when the user clicks the application icon 502 of the settings application, the mobile phone can respond to the user's operation and display the settings interface 503, as shown in Figure 5A(2). For example, the settings interface 503 may include one or more controls, including but not limited to: settings option 1, settings option 2, ..., settings option 7 and sound effect settings option 504, etc., which are not limited in this application.
[0154] Referring to Figure 5A(2), when a user needs to set the sound effect of the audio device playing the audio signal, they can click the sound effect setting option 504. The mobile phone can respond to the user's operation and display the sound effect setting interface 505, as shown in Figure 5B.
[0155] Figure 5B is a schematic diagram of another mobile phone interface according to an embodiment of this application.
[0156] Referring to FIG5B(1), the sound effect settings interface 505 may, by way of example, include one or more controls, including but not limited to: multiple scene options (such as movie scene options, sports scene options, game scene options, etc.), multiple environment options (such as gym options, park options, outdoor street options, etc.), and personalized sound effect settings options, etc., which are not limited in this application.
[0157] Referring to Figure 5B(1), when a user wants to set the current listening scene to a motion scene, they can click the motion scene option. At this time, the mobile phone can respond to the user's operation, mark the motion scene option, and obtain the scene input by the user as a motion scene. At this time, the listening scene can be determined to be a motion scene.
[0158] It should be understood that the listening scenario can also be determined based on at least two of the following: the type of the first audio signal, the device usage status, and the user input scenario; this can improve the accuracy of the determined listening environment.
[0159] For example, when the listening scenario is determined based on the type of the first audio signal, S401 can be executed first, followed by S402; when the listening scenario is determined based on the user-input scenario or device usage status, this application does not limit the execution order of S401 and S402.
[0160] S403, outputs a second audio signal and a third audio signal; wherein, the second audio signal is played through the air conduction audio playback unit, and the third audio signal is played through the bone conduction audio playback unit; when the first audio signal is the same, the second audio signal and the third audio signal output are different for different listening scenarios; when the listening scenario is the same, the second audio signal and the third audio signal output are different for different first audio signals.
[0161] For example, after acquiring the first audio signal and the listening scenario, the terminal device can perform audio processing on the first audio signal based on the listening scenario to generate a second audio signal and a third audio signal. For example, a corresponding audio effect link can be pre-set for each listening scenario; and a mapping relationship between each listening scenario and the corresponding audio effect link can be established, as shown in Table 1.
[0162] Table 1
[0163] Thus, after obtaining the listening scene, the corresponding audio effect link can be determined by looking up the mapping relationship; that is, configuring the audio effect link can be performed by the audio effect link configuration module 412 in Figure 3A. Subsequently, the audio effect processing module 414 in Figure 3A can perform audio effect processing on the first audio signal based on the audio effect link to generate the second and third audio signals.
[0164] Referring to Table 1, assuming the acquired listening scene is a motion scene, the corresponding audio effect link can be identified as audio effect link A1. In this case, audio effect processing can be performed on the first audio signal based on audio effect link A1 to generate the second and third audio signals. Assuming the acquired listening scene is a movie scene, the corresponding audio effect link can be identified as audio effect link A2. In this case, audio effect processing can be performed on the first audio signal based on audio effect link A2 to generate the second and third audio signals, and so on.
[0165] It can be seen that, when the first audio signal is the same, the second audio signal and the third audio signal output are different for different listening scenarios; when the listening scenario is the same, the second audio signal and the third audio signal output are different for different first audio signals.
[0166] Then, the terminal device can output the second audio signal to the air conduction audio playback unit of the audio device and output the third audio signal to the bone conduction audio playback unit; then, the air conduction audio playback unit plays the second audio signal and the bone conduction audio playback unit plays the third audio signal.
[0167] In some cases, the second audio signal can also be referred to as an air conduction signal, and in some cases, the third audio signal can also be referred to as a bone conduction signal.
[0168] It should be understood that when the air conduction audio playback unit and the bone conduction audio playback unit are not in the same audio device, the terminal device solves the communication delay between the terminal device and the air conduction audio playback unit, as well as the communication delay between the terminal device and the bone conduction audio playback unit, in the process of generating the second and third audio signals. Therefore, the synchronous playback of the second and third audio signals can be guaranteed.
[0169] In summary, this application takes into account the listening scenario and generates air conduction signals (i.e., the second audio signal) and bone conduction signals (i.e., the third audio signal) based on the listening scenario. When the listening scenario is different, the output air conduction signals and bone conduction signals are different; it can be seen that this application can adaptively adjust the output air conduction signals and bone conduction signals according to the listening scenario; in this way, the audio device can be applied to a variety of scenarios and can provide a relatively good listening experience in most scenarios.
[0170] For example, even with the same listening scenario, the user's expected listening experience will differ depending on the listening environment; therefore, this application can also consider the listening environment when generating air conduction signals and bone conduction signals to ensure that the audio device can provide a better listening experience.
[0171] Figure 6 is a schematic diagram of another audio processing procedure 600 according to an embodiment of this application.
[0172] S601, acquire the first audio signal, which is an audio signal from a multimedia file or a real-time call audio signal.
[0173] S602, obtain the listening scene.
[0174] For example, S601 to S602 can be referred to the description of S401 to S402 above, and will not be repeated here.
[0175] S603, obtain the listening environment.
[0176] For example, S603 can be performed by the environment recognition module 413 in FIG3A.
[0177] For example, the listening environment is the environment in which audio signals are played through audio devices. This listening environment is not limited to cinemas, elevators, shopping malls, subways, outdoor streets, mountains, living rooms, offices, etc.; it should be understood that the listening environment can be defined according to needs, and this application does not impose any restrictions on it.
[0178] In one possible implementation, the listening environment can be determined based on a fourth audio signal. This fourth audio signal is obtained by collecting ambient sound data and can also be referred to as an ambient sound audio signal. The fourth audio signal can be collected by a terminal device (including a terminal device executing the audio processing method of this application and / or other terminal devices (such as smartwatches)) or by an audio device. Exemplarily, the terminal device can perform environmental recognition based on the fourth audio signal to determine the listening environment. For example, the terminal device can extract key features from the fourth audio signal, such as the spectrum, Mel-frequency cepstral coefficients (MFCC), sound intensity, and frequency distribution. Subsequently, the extracted key features can be input into a classification model (which can be implemented by a pre-trained neural network), and the classification model outputs the listening environment.
[0179] In one possible implementation, the listening environment can be determined based on data from other sensors. This other sensor data can refer to data collected by sensors other than the microphone (such as Global Positioning System (GPS), accelerometer, gyroscope, barometer, light sensor, temperature sensor, magnetic field sensor, etc.). This other sensor data can be collected by the terminal device executing the audio processing method of this application and / or other terminal devices and / or audio devices. For example, the user's geographical location (e.g., living room, office, park) can be determined based on GPS data; subsequently, the listening environment can be determined based on the user's geographical location. Another example is determining the user's motion state (e.g., stationary, walking, running) based on data from an accelerometer; then, the listening environment (e.g., car, subway) can be determined based on the user's motion state. Yet another example is determining the rotation and orientation of the terminal device based on data from a gyroscope; then, the listening environment (e.g., elevator, stairs) can be determined based on the rotation and orientation of the terminal device. Finally, determining whether the user is in an indoor or outdoor environment can be determined based on data from a light sensor. For example, based on data collected by a magnetic field sensor, it can identify whether a user is near a specific device (such as a stereo or television); and further identify whether the user is in the living room or office.
[0180] One possible implementation is to determine the listening environment based on the device's usage status. For example, the listening environment can be determined based on the device's volume; high volume may indicate the user is in a noisy environment (such as an outdoor street or shopping mall), while low volume may indicate the user is in a quiet environment (such as a library or conference room). Alternatively, the listening environment can be determined based on the Wi-Fi information the device is connected to; for example, if the device is connected to office Wi-Fi, the listening environment is determined to be an office; if the device is connected to public Wi-Fi, the listening environment is determined to be a public place such as a coffee shop or airport. Furthermore, the listening scenario can be determined based on the application software being used on the device; for example, if video conferencing software is being used, the listening environment is determined to be an office; if navigation software is being used, the listening environment is determined to be outdoors.
[0181] In one possible implementation, the listening environment can be determined based on the user's input environment. Specifically, the user's input environment can be used to determine the listening environment.
[0182] Referring to Figure 5B(2), when a user wants to set the current listening environment to an outdoor street, they can click the outdoor street option. At this time, the mobile phone can respond to the user's operation, mark the outdoor street option, and obtain the environment input by the user as an outdoor street. At this time, it can be determined that the listening environment is an outdoor street.
[0183] It should be understood that this application can also determine the listening environment based on at least two of the following: a fourth audio signal, other sensor data, device usage status, or user-input environment; this can improve the accuracy of the determined listening environment.
[0184] It should be understood that this application does not restrict the execution order of S603, S602, and S601.
[0185] S604, output a second audio signal to the air conduction audio playback unit and output a third audio signal to the bone conduction audio playback unit; wherein, when the first audio signal is the same, the second audio signal and the third audio signal output are different for different listening scenarios; when the listening scenario is the same, the second audio signal and the third audio signal output are different for different first audio signals; when the first audio signal is the same and the listening scenario is the same, the second audio signal and the third audio signal output are different for different listening environments.
[0186] For example, after acquiring the first audio signal, the listening scenario, and the listening environment, the terminal device can perform audio processing on the first audio signal based on the listening scenario to generate a second audio signal and a third audio signal. For example, a mapping relationship between each set of listening environments and scenarios and the corresponding audio effect link can be pre-established; for example, the mapping relationship can be shown in Table 2 below:
[0187] Table 2
[0188] Thus, after obtaining the listening scene and listening environment, the corresponding audio effect link can be determined by looking up the mapping relationship; that is, the audio effect link can be configured by the audio effect link configuration module 412 in Figure 3A. Subsequently, the audio effect processing module 414 in Figure 3A can perform audio effect processing on the first audio signal based on the audio effect link to generate the second and third audio signals.
[0189] Referring to Table 2, assuming the acquired listening scenario is a sports scene and the listening environment is a park, the corresponding audio effect link can be identified as audio effect link B2. In this case, audio effect processing can be performed on the first audio signal based on audio effect link B2 to generate the second and third audio signals. Assuming the acquired listening scenario is a sports scene and the listening environment is a gym, the corresponding audio effect link can be identified as audio effect link B3. In this case, audio effect processing can be performed on the first audio signal based on audio effect link B3 to generate the second and third audio signals. Assuming the acquired listening scenario is a phone call scene and the listening environment is an elevator, the corresponding audio effect link can be identified as audio effect link B12. In this case, audio effect processing can be performed on the first audio signal based on audio effect link B12 to generate the second and third audio signals, and so on.
[0190] It should be understood that the listening scenario refers to the user's audio experience needs in a specific context; it focuses more on the user's behavior and purpose. The listening environment refers to the physical space in which the user is located and its acoustic characteristics; it focuses more on the impact of the external environment on the audio experience. Therefore, the listening scenario and listening environment together determine the sound effects processing method. By combining the listening scenario and listening environment, a more personalized and high-quality audio experience can be provided.
[0191] The process of generating the second and third audio signals will be explained below using Figure 3A as an example.
[0192] Figure 7A is a schematic diagram of another audio processing procedure 700 according to an embodiment of this application. The listening scene corresponding to the audio processing procedure 700 in Figure 7A is a motion scene, and the listening environment is an outdoor street.
[0193] For example, the scenario corresponding to Figure 7A can be as follows: When a user is about to run on an outdoor street, they will turn on the exercise recording function on their smartwatch, wear ear-hook headphones, and open a music app on their phone to listen to music. Then, they can start running on the outdoor street. During the run, the smartwatch will record the user's real-time heart rate, location, and other information to generate the running trajectory; and record the user's running cadence, pace, and other information; in addition, there may be environmental noise such as vehicles and pedestrians on the outdoor street. Among them, the bone conduction and air conduction two-in-one ear-hook headphones worn by the user can be as shown in Figure 2A(2); when worn, the headphones are hung on the ears, the air conduction audio playback unit is located in the concha cavity, and the sound is guided to the ear canal opening for playback through the sound guide hole, and the bone conduction audio playback unit is attached to the skin surface through the headphone shell and transmits the sound to the user through vibration.
[0194] S701, acquire the first audio signal.
[0195] For example, the first audio signal is the audio signal of music.
[0196] S702, obtain the listening scene.
[0197] For example, the scene recognition module 411 (which can use a traditional sound effect algorithm or a pre-trained neural network) detects that the smartwatch has started the motion recording function and the mobile phone has started the music app; at the same time, it detects the audio signal of the music; therefore, the scene recognition module 411 identifies the current listening scene as a motion scene.
[0198] S703, acquire the fourth audio signal.
[0199] S704, acquire other sensor data.
[0200] S705, obtain the listening environment.
[0201] For example, the environment recognition module 413 (which can use a traditional sound effect algorithm or a pre-trained neural network) detects that the phone's GPS location information is urban road, and the environmental noise collected by the headphones includes car horn sounds, outdoor street noise, etc.; therefore, the environment recognition module 413 identifies the current listening environment as an outdoor street.
[0202] For example, S701 to S705 can be described with reference to the above description, and will not be repeated here.
[0203] S706 determines the sound effect link based on the listening scenario and listening environment. This sound effect link includes the audio sound effect link and the environmental sound effect link.
[0204] For example, S706, which is configuring the audio link, can be executed by the audio link configuration module 412 in Figure 3A.
[0205] For example, the audio effect link may include an audio effect link and an ambient sound effect link; wherein, the audio effect link can be used to process the first audio signal, the ambient sound effect link can be used to process the fourth audio signal, and generate an audio signal for controlling the vibration intensity of the bone conduction audio playback unit. For example, the audio effect link may include a first sub-link and a second sub-link, the first sub-link being used to process all audio components contained in the first audio signal, and the second sub-link being used to process the rhythmic components in the first audio signal. For example, the ambient sound effect link may include a third sub-link and a fourth sub-link, wherein the third sub-link is used to process the fourth audio signal, and the fourth sub-link is used to generate an audio signal for controlling the vibration intensity of the bone conduction audio playback unit.
[0206] Assuming the scenario in Figure 7A is as described above, the listening scene is a motion scene, and the listening environment is an outdoor street. Therefore, according to Table 2, the audio effect link is audio effect link B1. The audio effect algorithms included in audio link B1 and their order can be shown in Figure 7B.
[0207] The first sub-link of the audio effects link in Figure 7B may include a dynamic adaptive frequency division algorithm, an air conduction audio DRC algorithm, a bone conduction audio DRC algorithm, an air conduction audio adaptive EQ algorithm, and a bone conduction audio adaptive EQ algorithm. The second sub-link of the audio effects link may include a low-frequency drum beat separation algorithm, a motion rhythm enhancement algorithm, and an audio enhancement algorithm.
[0208] The third sub-link of the ambient sound effect link in Figure 7B may include a key sound separation algorithm, an enhanced pass-through algorithm, an air-conducting audio ANC algorithm, and a bone-conducting audio ANC algorithm. The fourth sub-link of the ambient sound effect link may include a key sound separation algorithm, a vibration directionality simulation algorithm, and a vibration enhancement algorithm.
[0209] Then, the sound effects processing module 414 in Figure 3A executes the following steps S707 to S710:
[0210] S707 performs audio effect processing on the first audio signal based on the audio effect link to generate the first audio component in the second audio signal and the second audio component in the third audio signal.
[0211] For example, the audio processing module 414 can perform audio processing on the first audio signal according to the first sub-link to generate the first audio component in the second audio signal and the second audio component in the third audio signal.
[0212] Referring to Figure 7B, exemplarily, the sound effects processing module 414 can invoke a dynamic adaptive frequency division algorithm to perform frequency division processing on the first audio signal to obtain the high-frequency component (i.e., audio signal 1) and the low-frequency component (i.e., audio signal 2) in the first audio signal. Subsequently, the sound effects processing module 414 can invoke an air conduction audio DRC algorithm to perform dynamic range compression processing on audio signal 1 to obtain audio signal 3; and invoke a bone conduction audio DRC algorithm to perform dynamic range compression processing on audio signal 2 to obtain audio signal 4. Next, the sound effects processing module 414 can invoke an air conduction audio adaptive EQ algorithm to adjust the EQ of audio signal 3 to obtain the first audio component in the second audio signal; and can invoke a bone conduction audio adaptive EQ algorithm to adjust the EQ of audio signal 4 to obtain the second audio component in the third audio signal.
[0213] S708 performs audio effect processing on the fourth audio signal based on the ambient sound effect link to generate the third audio component in the second audio signal, the fourth audio component in the third audio signal, and the sixth audio component in the third audio signal.
[0214] For example, the audio processing module 414 can perform audio processing on the fourth audio signal according to the third sub-link to generate the third audio component in the second audio signal and the fourth audio component in the third audio signal; and perform audio processing on the fourth audio signal according to the fourth sub-link to generate the sixth audio component in the third audio signal.
[0215] Referring to Figure 7B, exemplarily, the audio processing module 414 can invoke a key sound separation algorithm to separate key sounds (including key human voices and key ambient sounds) from the fourth audio signal to obtain audio signal 7. Key ambient sounds include key sounds in the current environment that may affect the user, such as car horns, traffic light signals, and alarms in an outdoor street environment. Key human voices include voiceprints of friends in the user's address book, user-defined keywords such as the user's name, and the voice of a person speaking to the user. Subsequently, the audio processing module 414 can invoke an enhancement pass-through algorithm to enhance audio signal 7 (e.g., increase loudness, enhance signal-to-noise ratio, etc.) to obtain audio signal 8. These enhanced key ambient sounds or key human voices will also be passed through in the form of spatial audio, allowing the user to quickly locate the real-world sound source. Next, the sound effects processing module 414 can call the air conduction audio ANC algorithm to actively reduce noise in the audio signal 8 to obtain the third audio component in the second audio signal; and call the bone conduction audio ANC algorithm to actively reduce noise in the audio signal 8 to obtain the fourth audio component in the third audio signal. In this way, other sounds in the environment can be filtered out.
[0216] Referring again to Figure 7B, exemplarily, the sound effect processing module 414 can invoke a vibration directionality simulation algorithm to process the audio signal 7, generating an audio signal 10 that characterizes the vibration intensity of the bone conduction audio playback unit (audio signal 10 may include two audio signals, corresponding to the left and right earphones respectively; these two audio signals have different loudnesses). Subsequently, the sound effect processing module 414 can invoke a vibration enhancement algorithm to process the audio signal 10 (for example, it can divide the two audio signals in the audio signal 10 by frequency, and separate the audio frequencies near the f0 frequency of the bone conduction audio playback unit in the two audio signals and increase their loudness), to generate an audio signal that can control the vibration intensity of the bone conduction audio playback unit, namely the sixth audio component in the third audio signal. In this way, by simulating the different strengths of the bone conduction vibrations of the left and right earphones, directionality is simulated, indicating the location of the sound source to the user.
[0217] For example, when a user runs across an intersection, a car approaches from the left and honks its horn. The headphones identify the horn sound and transmit it to the user, amplifying the playback. Using technologies such as beamforming, the headphones pinpoint the source of the horn and render the transmitted sound as originating from the left. Simultaneously, the left bone conduction oscillator in the bone conduction audio playback unit vibrates significantly stronger than the right bone conduction oscillator, indicating to the user that the horn is honking from the left.
[0218] S709 performs sound effect processing on the rhythm component in the first audio signal based on the audio effect link and other sensor data to generate the fifth audio component in the second audio signal and the tenth audio component in the third audio signal.
[0219] For example, the sound effects processing module 414 can call the second sub-link and combine it with other sensor data (such as motion sensor data) to perform sound effects processing on the rhythm component in the first audio signal to generate the fifth audio component in the second audio signal and the tenth audio component in the third audio signal.
[0220] Referring again to Figure 7B, exemplarily, the sound effect processing module 414 can invoke a low-frequency bone conduction separation algorithm to separate the audio component (a type of rhythm component) corresponding to the low-frequency drum beats from the first audio signal to obtain audio signal 5. Subsequently, the sound effect processing module 414 can invoke a motion rhythm enhancement algorithm to perform motion rhythm enhancement processing on audio signal 5 based on motion sensor data (e.g., retaining the drum beat segments in audio signal 5 that have the same frequency as the step frequency in the motion sensor data, and muting the others) to obtain audio signal 6. Next, the sound effect processing module 414 can invoke an audio enhancement algorithm to perform audio enhancement processing on audio signal 6 (e.g., periodically enhancing the loudness of drum beats with the same frequency as the step frequency) to obtain the fifth audio component in the second audio signal and the tenth audio component in the third audio signal.
[0221] S710 performs sound effect processing on the rhythm component in the first audio signal based on the audio effect link, the ambient sound effect link, and other sensor data to generate the eighth audio component in the third audio signal.
[0222] For example, the sound effect processing module 414 can call some sound effect algorithms in the second sub-link and some sound effect algorithms in the fourth sub-link, and combine them with other sensor data to perform sound effect processing on the rhythm component in the first audio signal to generate the eighth audio component in the third audio signal.
[0223] Referring again to Figure 7B, exemplarily, the audio processing module 414 can call the low-frequency bone conduction separation algorithm and the dynamic rhythm enhancement algorithm in the second sub-link to process the first audio signal and generate audio signal 6. Then, it can call the vibration directionality module algorithm in the fourth sub-link to process audio signal 6 to obtain audio signal 9. Subsequently, the audio processing module 414 can call the vibration enhancement algorithm to process audio signal 9 to generate an audio signal that can control the vibration intensity of the bone conduction audio playback unit, i.e., the eighth audio component in the third audio signal.
[0224] It should be noted that the sixth and eighth audio components in the third audio signal are different from the second and fourth audio components in the third audio signal. The sixth and eighth audio components are audio signals used to control the vibration intensity of the bone conduction audio playback unit so that the user can feel the vibration; while the second and fourth audio components are audio signals used to control the vibration of the bone conduction audio playback unit to generate sound so that the user can feel the sound.
[0225] Furthermore, when the smartwatch detects that the user's heart rate exceeds a preset heart rate threshold, or that the user's cadence is below a preset cadence threshold, the smartwatch can send a signal to the terminal device. At this time, the terminal device can also generate a first audio component and / or a second audio component via a first sub-link, and play them through the air conduction audio playback unit and / or bone conduction audio playback unit in the headphones, to alert the user that their current exercise or physical state has exceeded the set threshold. Alternatively, the terminal device can also generate an eighth audio component by calling a vibration enhancement algorithm and play it through the bone conduction audio playback unit; thereby helping to improve the user's exercise performance.
[0226] S711 outputs a second audio signal containing a first audio component, a third audio component, and a fifth audio component to the air-conducting audio playback unit.
[0227] S712 outputs a third audio signal containing the second, fourth, sixth, eighth, and tenth audio components to the bone conduction audio playback unit.
[0228] Figure 8A is a schematic diagram of another audio processing procedure 800 according to an embodiment of this application. The listening scenario corresponding to the audio processing procedure 800 in Figure 8A is a call scenario, and the listening environment is an elevator.
[0229] For example, the scenario corresponding to Figure 8A can be as follows: The user wears smart glasses (containing an air conduction audio playback unit) and bone conduction headphones (containing a bone conduction audio playback unit) to listen to music while working in the office. After get off work, the user takes the elevator and there are other people riding the elevator with the user. After receiving a call, the user answers using bone conduction headphones + smart glasses (as shown in Figure 2B(2)). During the call, the voice of the other user is guided to the ear canal through the sound guide hole contained in the air conduction audio playback unit of the smart glasses, and the sound is transmitted to the user through the vibration of the bone conduction audio playback unit of the bone conduction headphones.
[0230] S801, acquire the first audio signal.
[0231] For example, the first audio signal is the voice of the user on the other end of the call and the audio signal of the environment where the user is located.
[0232] S802, obtain the listening scene.
[0233] For example, the scene recognition module 411 detects that the mobile phone is currently in a call state, and also recognizes that the first audio signal is the audio signal of a human voice call; therefore, the scene recognition module 411 can identify the current listening scene as a call scene.
[0234] S803, acquire the fourth audio signal.
[0235] S804, obtain the listening environment.
[0236] For example, the environment recognition module 413 detects that the gyroscope information of the mobile phone shows vertical up and down movement and that the mobile phone signal is weak; the environmental noise collected by the headphones contains significant reverberation and human voices, etc.; therefore, the environment recognition module 413 identifies the current environment as an elevator.
[0237] For example, S801 to S804 can be described with reference to the above description, and will not be repeated here.
[0238] S805 determines the sound effect link based on the listening scenario and listening environment. This sound effect link includes the audio sound effect link and the environmental sound effect link.
[0239] For example, S805, which configures the audio effect link, can be executed by the audio effect link configuration module 412 in Figure 3A. For example, S805 can be described with reference to S706 above, and will not be repeated here. The difference between the audio effect link in S805 and the audio effect link in S706 is that the ambient sound effect link in S805 only includes a third sub-audio effect link, which is used to process the fourth audio signal; the audio effect link in S805 includes a first sub-link and a second sub-link, which are used to perform audio effect processing on different audio components in the first audio signal.
[0240] Assuming the scenario in Figure 8A is as described above, the listening scenario is a phone call scenario, and the listening environment is an elevator. Therefore, according to Table 2, the audio effect link is audio effect link B12. The audio effect algorithms included in audio link B12 and their order can be shown in Figure 8B.
[0241] The first sub-link of the audio effects link in Figure 8B may include a voice separation algorithm, a sound leakage suppression algorithm, a voice enhancement algorithm, and a vibration suppression algorithm. The second sub-link of the audio effects link may include a background sound separation algorithm, a noise reduction algorithm, and an audio enhancement algorithm.
[0242] The third sub-link in the environmental sound effect link in Figure 8B may include the air conduction audio ANC algorithm and the bone conduction audio ANC algorithm.
[0243] Then, the sound effects processing module 414 in Figure 3A executes the following steps S806 to S810:
[0244] S806 performs audio effect processing on the first audio signal based on the audio effect link to generate the first audio component in the second audio signal and the second audio component in the third audio signal.
[0245] For example, the audio processing module 414 can perform audio processing on the first audio signal according to the first sub-link to generate the first audio component in the second audio signal and the second audio component in the third audio signal.
[0246] Referring to Figure 8B, exemplarily, the audio processing module 414 can call a voice separation algorithm to separate the audio component corresponding to the human voice from the first audio signal to obtain audio signal 1. Subsequently, the audio processing module 414 can call a sound leakage suppression algorithm to perform sound leakage suppression processing on audio signal 1 to obtain the first audio component in the second audio signal; thereby preventing sound leakage from the smart glasses' speaker. The audio processing module 414 can also call a voice enhancement algorithm to perform voice enhancement processing on audio signal 1 (such as increasing loudness, increasing signal-to-noise ratio, etc.) to obtain audio signal 2; this can increase the proportion of speech in the bone conduction channel. Next, the audio processing module 414 can call a vibration suppression algorithm to perform vibration suppression processing on audio signal 2 to obtain the second audio component in the third audio signal; this can suppress low frequencies in the human voice, ensuring the comfort of the bone conduction audio playback unit and reducing sound leakage.
[0247] In addition, the sound processing module 414 can also call the sound leakage suppression algorithm to generate an audio signal (which is part of the first audio component) that is out of phase with the sound leakage of the bone conduction audio playback unit collected by the sound playback monitoring module, so as to cancel the sound leakage of the bone conduction headphones and achieve the purpose of active sound leakage compensation.
[0248] S807 performs audio processing on the fourth audio signal based on the ambient sound effect link to generate the third audio component in the second audio signal and the fourth audio component in the third audio signal.
[0249] For example, the audio processing module 414 can perform audio processing on the fourth audio signal according to the third sub-link to generate the third audio component in the second audio signal and the fourth audio component in the third audio signal.
[0250] Referring to Figure 8B, exemplarily, the sound processing module 414 can invoke the air conduction audio ANC algorithm to actively reduce noise in the fourth audio signal to obtain the third audio component in the second audio signal; and invoke the bone conduction audio ANC algorithm to actively reduce noise in the fourth audio signal to obtain the fourth audio component in the third audio signal. This allows other sounds in the environment to be filtered out.
[0251] Optionally, the environmental sound effect link in Figure 8B may also include an enhanced pass-through algorithm; after the enhanced pass-through algorithm is called to process the fourth audio signal, the air conduction audio ANC algorithm and the bone conduction audio ANC algorithm are called respectively to process the enhanced pass-through fourth audio signal, so that the user can quickly locate the sound source in the environment.
[0252] S808 performs audio effect processing on the first audio signal based on the audio effect link to generate the fifth audio component in the second audio signal and the tenth audio component in the third audio signal.
[0253] For example, the audio processing module 414 can call the second sub-link to perform audio processing on the first audio signal to generate the fifth audio component in the second audio signal and the tenth audio component in the third audio signal.
[0254] Referring again to Figure 8B, exemplarily, the audio effects processing module 414 can invoke a background sound separation algorithm to separate the audio components corresponding to the background sound from the first audio signal to obtain audio signal 3. Subsequently, the audio effects processing module 414 can invoke a noise reduction algorithm to perform noise reduction processing on audio signal 5 to obtain audio signal 4. Next, the audio effects processing module 414 can invoke an audio enhancement algorithm to perform audio enhancement processing on audio signal 4 to obtain the fifth audio component in the second audio signal and the tenth audio component in the third audio signal.
[0255] S809 outputs a second audio signal containing a first audio component, a third audio component, and a fifth audio component to the air-conducting audio playback unit.
[0256] S810 outputs a third audio signal containing the second, fourth, and tenth audio components to the bone conduction audio playback unit.
[0257] Figure 9A is a schematic diagram of another audio processing procedure 900 according to an embodiment of this application. The listening scenario corresponding to the audio processing procedure 800 in Figure 9A is a conference scene, and the listening environment is a conference room.
[0258] [Corrected according to Rule 91 03.11.2025] For example, the scenario corresponding to Figure 9A can be as follows: When a user is working in the office, they use an open-back air-conduction headset alone. When an online meeting is required and there are participants speaking less common languages, the user can use an air-conduction headset with a simultaneous interpretation module (as shown in Figure 2B(3)) to participate in the meeting. The user can hang the ear hook on their ear, place the air-conduction headset in the concha or ear canal, and place the simultaneous interpretation module behind the auricle, in front of the tragus, or around the ear; this ensures that the simultaneous interpretation module and the air-conduction headset are fixed on the ear and remain still. During the meeting, the voices of other participants are guided to the ear canal opening through the sound guide hole included in the air-conduction audio playback unit of the air-conduction headset, and the sound is transmitted to the user through the vibration of the bone conduction audio playback unit of the simultaneous interpretation module.
[0259] Figures 9B and 9C are schematic diagrams of the mobile phone interface.
[0260] For example, during the connection process between the air-conduction headset and the simultaneous interpretation module, the mobile phone can display the extension module settings interface 90, as shown in Figure 9B(1). The extension module settings interface 90 shows a schematic of the connection status between the air-conduction headset and the extension module, that is, the air-conduction headset is not yet communicating with the extension module and is searching for the extension module. After the air-conduction headset and the simultaneous interpretation module are connected, the extension module settings interface 90 can display a schematic of the headset and the simultaneous interpretation module being connected, as shown in Figure 9B(2). In addition, after the air-conduction headset and the simultaneous interpretation module are connected, the extension module settings interface 90 can also display the listening scene, the prompt message that simultaneous interpretation is enabled, and the simultaneous interpretation settings options.
[0261] For example, when a user clicks the simultaneous interpretation settings option in Figure 9B(2), the mobile phone can respond to the user's operation and display the simultaneous interpretation settings interface 91, as shown in Figure 9C. The simultaneous interpretation settings interface 91 includes a diagram showing that the headset and the simultaneous interpretation module are connected, as well as language options such as English, French, German, Spanish, and Russian. Users can click the corresponding language option as needed.
[0262] S901, acquire the first audio signal.
[0263] For example, the first audio signal is the audio signal of the conference audio.
[0264] S902, obtain the listening scene.
[0265] For example, after the scene recognition module 411 detects that the headphones are connected to the simultaneous interpretation module, it can identify the current listening scene as a conference scene. In addition, the scene recognition module 411 can further identify the scene as a **(language) conference scene based on the language option clicked by the user in the simultaneous interpretation settings interface 91, or based on the language detected by the audio signal collected by the microphone of the headphones or the audio signal played by the headphones.
[0266] S903, acquire the fourth audio signal.
[0267] S904, obtain the listening environment.
[0268] For example, the environment recognition module 413 identifies the human voice as clear, with indoor reverberation and low background noise based on the audio signal collected by the microphone of the headphones, and the GPS information in the mobile phone indicates that it is an office and the listening scene is a meeting scene; therefore, the environment recognition module 413 identifies the current environment as a meeting room.
[0269] For example, S901 to S904 can be described with reference to the above description, and will not be repeated here.
[0270] S905 determines the sound effect link based on the listening scenario and listening environment. This sound effect link includes the audio sound effect link and the environmental sound effect link.
[0271] For example, S905, which configures the audio effect link, can be executed by the audio effect link configuration module 412 in Figure 3A. For example, S905 can be described with reference to S706 above, and will not be repeated here. The difference between the audio effect link in S905 and the audio effect link in S706 is that the ambient sound effect link in S905 only includes a third sub-audio effect link, which is used to process the fourth audio signal; the audio effect link in S905 only includes a first sub-link, which is used to process the first audio signal.
[0272] Assuming the scenario corresponding to Figure 9A is as shown above, the listening scenario is a conference scenario, and the listening environment is a conference room; then, according to Table 2, the audio effect link can be determined to be audio effect link B22. The audio effect algorithms included in audio link B22 and the order of each audio effect algorithm can be shown in Figure 9D.
[0273] The first sub-link of the audio effect link in Figure 9D may include a voice separation algorithm, an ANC algorithm, a sound leakage suppression algorithm, and a vibration suppression algorithm. The third sub-link of the ambient sound effect link in Figure 9D may include an enhanced pass-through algorithm, an air conduction audio ANC algorithm, and a bone conduction audio ANC algorithm.
[0274] Then, the sound effects processing module 414 in Figure 3A executes the following S906 to S907:
[0275] S906, based on the audio effect link, performs sound effect processing on the first audio signal to generate the first audio component in the second audio signal and the second audio component in the third audio signal.
[0276] For example, the audio processing module 414 can perform audio processing on the first audio signal according to the first sub-link to generate the first audio component in the second audio signal and the second audio component in the third audio signal.
[0277] Referring to Figure 9D, exemplarily, the audio processing module 414 can call a voice separation algorithm to separate the audio component corresponding to the human voice from the first audio signal to obtain audio signal 1. Subsequently, the audio processing module 414 can call an ANC algorithm to suppress noise in audio signal 1 to obtain audio signal 2, which is then output to the translation module. The translation module can perform speech recognition on audio signal 2 to obtain first text; then, the translation module can translate the first text into second text according to the language selected by the user or the system language of the terminal device; subsequently, it uses text-to-speech (TTS) technology to convert the second text into audio signal 3. Next, the audio processing module 414 can call a sound leakage suppression algorithm to suppress sound leakage in audio signal 3 to obtain the first audio component in the second audio signal; thereby preventing sound leakage from the speaker of the air-conduction headphones. The sound effects processing module 414 can also call the vibration suppression algorithm to perform vibration suppression processing on the audio signal 3 to obtain the second audio component in the third audio signal; in this way, the low frequency in the human voice can be suppressed, ensuring the comfort of the bone conduction audio playback unit and reducing sound leakage.
[0278] In addition, the sound processing module 414 can also call the sound leakage suppression algorithm to generate an audio signal (part of the first audio component) that is out of phase with the sound leakage of the bone conduction audio playback unit collected by the sound playback monitoring module, so as to cancel the sound leakage of the simultaneous interpretation module and achieve the purpose of active sound leakage compensation.
[0279] S907 performs audio processing on the fourth audio signal based on the ambient sound effect link to generate the third audio component in the second audio signal and the fourth audio component in the third audio signal.
[0280] For example, the audio processing module 414 can perform audio processing on the fourth audio signal according to the third sub-link to generate the third audio component in the second audio signal and the fourth audio component in the third audio signal.
[0281] Referring to Figure 9B, exemplarily, the audio processing module 414 can invoke an enhanced pass-through algorithm to perform enhanced pass-through processing on the fourth audio signal to obtain audio signal 4. Then, the audio processing module 114 can invoke an air conduction audio ANC algorithm to perform active noise reduction processing on the audio signal 4 to obtain the third audio component in the second audio signal; and invoke a bone conduction audio ANC algorithm to perform active noise reduction processing on the audio signal 4 to obtain the fourth audio component in the third audio signal.
[0282] S908 outputs a second audio signal containing a first audio component and a third audio component to the air-conducting audio playback unit.
[0283] S909 outputs a third audio signal containing a second audio component and a fourth audio component to the bone conduction audio playback unit.
[0284] It should be understood that if the listening scenario is different from the listening scenario shown in the above embodiments, or if the listening environment is different from the listening environment shown in the above embodiments, the corresponding sound effect link may also be different from the sound effect link shown in the above embodiments. These will not be illustrated here.
[0285] In summary, this application can convert the first audio signal into air-conducted audio components, bone-conducted audio components, and bone-conducted vibration components. The bone-conducted vibration component controls the vibration of the bone-conducting oscillator, thus providing vibrational cues to the user. It can also indicate the spatial location of key ambient sounds. Furthermore, it adaptively switches between matching audio effects in different listening scenarios and environments to adjust the proportion and perceived quality of each component, ensuring optimal sound quality and the best overall experience. This guarantees a consistent listening experience across multiple listening scenarios and environments. In addition, this application also transmits key ambient sounds and filters out noise, indicating the location of ambient sound sources to the user.
[0286] For example, users can also customize the audio effect chain configuration; in this way, users' personalized listening requirements can be met.
[0287] Figures 10A to 10E are schematic diagrams of mobile phone interfaces.
[0288] For example, after the mobile phone identifies the listening environment, it can switch the listening environment of the audio device to that listening environment and display a prompt message, as shown in Figure 10A. If the audio processing process 700 identifies the listening environment as an outdoor street, the prompt message displayed is shown in Figure 10A(1); if the audio processing process 800 identifies the listening environment as an elevator, the prompt message displayed is shown in Figure 10A(2); if the audio processing process 900 identifies the listening environment as a conference room, the prompt message displayed is shown in Figure 10A(3).
[0289] In one possible implementation, when a user wants to set the sound effects of the audio device, they can click the settings option in Figure 10A; the mobile phone can respond to the user's operation and display the personalized sound effects settings interface 1001, as shown in Figure 10C.
[0290] In one possible implementation, when a user wishes to adjust the sound effects of an audio device, they can click on the personalized sound effect settings option in the sound effect settings interface 505. The phone responds to the user's action by displaying the personalized sound effect settings interface 1001, as shown in Figure 10C. The personalized sound effect settings interface 1001 may include one or more controls, including but not limited to: frequency response settings for the left air conduction channel, frequency response settings for the left bone conduction channel, frequency response settings for the right air conduction channel, frequency response settings for the right bone conduction channel, vibration enhancement settings, and sound effect enhancement settings. The left air conduction channel may refer to the left (left ear) air conduction audio playback unit, the right air conduction channel may refer to the right (right ear) air conduction audio playback unit, the left bone conduction channel may refer to the left (left ear) bone conduction audio playback unit, and the right bone conduction channel may refer to the right (right ear) bone conduction audio playback unit.
[0291] Referring again to Figure 10C, when the user clicks the frequency response setting option for the left bone conduction channel, the phone responds to the user's action, marks the frequency response setting option for the left bone conduction channel, and displays the corresponding frequency response curve setting option 1002. The frequency response curve setting option 1002 includes four sliders: "Very Low Frequency," "Low Frequency," "Mid Frequency," and "High Frequency." The user can set the frequency response curve corresponding to the left bone conduction channel by sliding any of the sliders.
[0292] Referring to Figure 10C, users can tap the toggle switch for the vibration enhancement option, and the phone can respond to the user's action by turning the vibration enhancement function on or off. Similarly, users can tap the toggle switch for the sound enhancement option, and the phone can respond to the user's action by turning the sound enhancement function on or off.
[0293] Based on the user's settings in the personalized sound effects settings interface 1001, the mobile phone can obtain the sound effect information input by the user. Then, in one possible implementation, the mobile phone can obtain the sound effect path based on the listening scenario and the user-input sound effect information; or, it can obtain the sound effect path based on the listening scenario, listening environment, and user-input sound effect information. In another possible implementation, the mobile phone can adjust the sound effect path determined based on the listening scenario based on the user-input sound effect information; or, it can adjust the sound effect path determined based on the listening scenario and listening environment based on the user-input sound effect information.
[0294] Referring again to Figure 10C, when the user clicks the personalized sound effect settings option in the sound effect settings interface 505, the phone responds to the user's operation and displays the mode selection interface 1002, as shown in Figure 10D. The mode selection interface 1002 may include, but is not limited to, professional mode options, normal mode options, etc. When the user clicks the normal mode option, the phone can respond to the user's operation and display the personalized sound effect settings interface 1001, as shown in Figure 10C. When the user clicks the professional mode option, the phone can respond to the user's operation and display the personalized sound effect settings interface 1003, as shown in Figure 10E(1). The personalized sound effect settings interface 1003 may include one or more controls, including but not limited to: listening environment information (such as sports scene), listening environment information (such as outdoor street), sound effect link settings area, etc. Among them, the sound effect link settings area may include multiple sound effect algorithm settings options: sound effect algorithm 1 settings option, sound effect algorithm 2 settings option, sound effect algorithm 3 settings option, sound effect algorithm 4 settings option... Each sound effect algorithm settings option may include deletion options and adjustment options. The sound effect algorithm setting area includes sound effect algorithm setting options that correspond one-to-one with the sound effect algorithms included in the sound effect link (referred to as sound effect link 1 for ease of description) determined by the mobile phone based on the listening scene (movement scene in Figure 10E(1)) and listening environment (outdoor street in Figure 10E(1)).
[0295] Referring to Figure 10E(1), when the user clicks the delete option in the sound effect algorithm 1 settings, the phone responds to the user's operation and deletes sound effect algorithm 1 in sound effect link 1. When the user clicks the management option in the sound effect algorithm 1 settings, the phone responds to the user's operation and displays the sound effect algorithm parameter adjustment interface 1004, as shown in Figure 10E(2). The sound effect algorithm parameter adjustment interface 1004 may include one or more controls, including but not limited to: parameter 1 setting option, parameter 2 setting option, parameter 3 setting option, parameter 4 setting option, etc. The user can enter parameter values in the parameter 1 setting option of sound effect algorithm 1 or click "-" or "+" to adjust the parameter values; in this way, the phone can obtain the parameter values (a kind of sound effect information) of parameter 1 included in sound effect algorithm 1 entered by the user, and then replace the parameter values of parameter 1 included in sound effect algorithm 1 in sound effect link 1 with the parameter values of parameter 1 included in sound effect algorithm 1 entered by the user. The same applies to other parameters, which will not be described in detail here.
[0296] In one example, FIG11 shows a schematic block diagram of an apparatus 1100 according to an embodiment of the present application. The apparatus 1100 may include a processor 1101 and a transceiver 1102, and optionally, a memory 1103.
[0297] The various components of device 1100 are coupled together via bus 1104, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as bus 1104 in the figure.
[0298] Optionally, the memory 1103 can be used to store instructions from the foregoing method embodiments. The processor 1101 can be used to execute the instructions in the memory 1103, control the transceiver 1102 to receive signals, and control the transceiver 1102 to transmit signals.
[0299] The device 1100 may be an electronic device or a chip of an electronic device in the above method embodiments. For example, the electronic device may be a terminal device, an audio device, etc.
[0300] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0301] This application also provides a chip, including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the above-described related method are executed to achieve the method in the above embodiments. The interface circuit is a transceiver 1102.
[0302] This embodiment also provides a computer-readable storage medium storing computer instructions. When these computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the methods described in the above embodiments. Exemplarily, the computer-readable storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0303] This embodiment also provides a computer program product containing computer instructions that, when executed by a computer or processor, cause the computer to perform the aforementioned steps to implement the methods described in the above embodiments. Exemplarily, the computer program product can be stored in random access memory (RAM), flash memory, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, read-only optical discs (CD-ROMs), or any other form of storage medium known in the art.
[0304] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0305] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0306] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0307] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0308] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.
[0309] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An audio processing method, characterized in that, The method includes: Acquire a first audio signal, wherein the first audio signal is an audio signal in a multimedia file or a communication audio signal; Obtain the listening scenario; Output the second and third audio signals; The second audio signal is played through the air conduction audio playback unit, and the third audio signal is played through the bone conduction audio playback unit. When the first audio signal is the same, the second and third audio signals output for different listening scenarios are different. When the listening scenarios are the same, the second and third audio signals output for different first audio signals are different.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the listening environment; Wherein, when the first audio signal is the same and the listening scenario is the same, the second and third audio signals output for different listening environments are different.
3. The method according to claim 1 or 2, characterized in that, The third audio signal includes bone conduction audio components and bone conduction vibration components.
4. The method according to any one of claims 1 to 3, characterized in that, The second audio signal includes a first audio component and a third audio component, and the third audio signal includes a second audio component and a fourth audio component, wherein the second audio component and the fourth audio component are bone conduction audio components; the method further includes: Acquire a fourth audio signal, which is obtained by collecting ambient sound. The first audio component and the second audio component are generated based on the first audio signal, and the third audio component and the fourth audio component are generated based on the fourth audio signal.
5. The method according to claim 4, characterized in that, The third audio signal also includes a sixth audio component, which is a bone conduction vibration component and is generated based on the fourth audio signal.
6. The method according to claim 4 or 5, characterized in that, The third audio signal also includes an eighth audio component, which is a bone conduction vibration component and is generated based on the rhythm component in the first audio signal.
7. The method according to claim 1 or 2, characterized in that, The method further includes: The sound effect link is obtained based on the listening scenario, or the sound effect link is obtained based on the listening scenario and the sound effect information input by the user. The first audio signal is processed according to the audio effect link to obtain the second audio signal and the third audio signal.
8. The method according to claim 2, characterized in that, The method further includes: The sound effect link is obtained based on the listening scenario and the listening environment, or the sound effect link is obtained based on the listening scenario, the listening environment, and the sound effect information input by the user. The first audio signal is processed according to the audio effect link to obtain the second audio signal and the third audio signal.
9. The method according to any one of claims 4 to 6, characterized in that, The method further includes: The first audio signal is processed based on the audio effect links contained in the audio effect link to obtain the first audio component and the third audio component; The fourth audio signal is processed based on the ambient sound effect link included in the sound effect link to obtain the second audio component and the fourth audio component; The sound effect link is determined based on the listening scenario and the listening environment, or the audio link is determined based on the listening scenario, the listening environment, and the sound effect information input by the user.
10. The method according to any one of claims 1 to 9, characterized in that, The acquisition of the listening scenario includes: The listening scenario is determined based on at least one of the user-input scenario, the type of the first audio signal, or the device usage status.
11. The method according to claim 2, characterized in that, The acquisition of the listening environment includes: The listening environment is determined based on at least one of the fourth audio signal, other sensor data, device usage status, or user-input environment. The fourth audio signal is obtained by collecting ambient sound.
12. The method according to any one of claims 4 to 6, characterized in that, The second audio signal also includes a fifth audio component, which is generated based on the rhythm component in the first audio signal.
13. An audio system, characterized in that, The audio system includes audio equipment and terminal equipment, and the audio equipment includes an air conduction audio playback unit and a bone conduction audio playback unit. The terminal device is configured to acquire a first audio signal, which is an audio signal in a multimedia file or a communication audio signal; acquire a listening scenario; output a second audio signal to the air conduction audio playback unit; and output a third audio signal to the bone conduction audio playback unit. Wherein, when the first audio signal is the same, the second and third audio signals output for different listening scenarios are different; when the listening scenarios are the same, the second and third audio signals output for different first audio signals are different.
14. The audio system according to claim 13, characterized in that, The audio device includes headphones and an expansion module, the expansion module including at least one of an audio module, a health monitoring module, or an environmental monitoring module; The headphones include the air conduction audio playback unit, and the expansion module includes the bone conduction audio playback unit.
15. An audio device, characterized in that, The audio device includes an air conduction audio playback unit and a bone conduction audio playback unit. The audio device is used to: acquire a first audio signal, wherein the first audio signal is an audio signal in a multimedia file or a communication audio signal; acquire a listening scenario; output a second audio signal to the air conduction audio playback unit; and output a third audio signal to the bone conduction audio playback unit. Wherein, when the first audio signal is the same, the second and third audio signals output for different listening scenarios are different; when the listening scenarios are the same, the second and third audio signals output for different first audio signals are different.
16. The audio device according to claim 15, characterized in that, The audio device includes headphones and an expansion module, the expansion module including at least one of an audio module, a health monitoring module, or an environmental monitoring module; The headphones include the air conduction audio playback unit, and the expansion module includes the bone conduction audio playback unit.
17. An electronic device, characterized in that, include: A memory and a processor, wherein the memory is coupled to the processor; The memory stores program instructions that, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1 to 12.
18. A chip, characterized in that, It includes one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method as described in any one of claims 1 to 12 are performed.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed on a computer or processor, causes the computer or processor to perform the method as described in any one of claims 1 to 12.
20. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a computer or processor, cause the steps of the method as described in any one of claims 1 to 12 to be performed.