Systems, methods, and apparatuses

WO2026072975A3PCT designated stage Publication Date: 2026-05-07DOLBY LABORATORIES LICENSING CORP
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2025-09-26
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Wearable devices face challenges in maintaining consistent media playback and user experience in noisy environments due to auditory masking effects, vocal sound interference, and environmental sound fluctuations, which affect loudness, timbre, and speech intelligibility.

Method used

A context-aware noise compensation system using sensors like microphones and accelerometers to detect vocal and environmental sounds, optimizing media playback by adjusting gain based on context information, and selectively passing important sound events.

Benefits of technology

Enhances media playback consistency and user experience by minimizing noise fluctuations and allowing selective passing of important sound events, improving loudness and speech clarity in noisy conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025048208_07052026_PF_FP_ABST
    Figure US2025048208_07052026_PF_FP_ABST
Patent Text Reader

Abstract

In one possible example aspect, the present disclosure provides a context-aware noise compensation system for a wearable device playing back media content to a user in a noisy environment, wherein the system is configured to: obtain a sound signal from the noisy environment; determine context information associated with the sound signal, wherein the context information comprises at least one of: vocal sound detection information indicative of presence of one or more vocal sounds of the user in the sound signal, or sound event detection information indicative of presence of one or more environmental sound events in the sound signal; and optimize, based on the context information, the media content for playback by the wearable device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]SYSTEMS, METHODS, AND APPARATUSES CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application Serial No.63 / 888,151, filed September 25, 2025, and claims priority to PCT / CN2024 / 122008, filed September 27, 2024, and claims priority to PCT / CN2024 / 121951, filed September 27, 2024, and claims priority to PCT / CN2024 / 121969, filed September 27, 2024, each of which is hereby incorporated by reference in its entirety. TECHNICAL FIELD The present disclosure generally relates to the technical field of signal processing, and more particularly, to various audio signal processing techniques for use in mobile and / or wearable devices. BACKGROUND Nowadays, wearable devices (or, more generally, mobile devices), such as headphones, earbuds, smart glasses, and head-mounted devices (HMDs), etc., are increasingly used in noisy and variable environments. In a broad sense, users / customers (e.g., wearers of such mobile / wearable devices) may typically expect consistent media playback, awareness of relevant ambient sounds (or sound events), natural conversational support, or the like, which may pose various challenges or issues. For instance, in some possible use cases / scenarios, media playback by various wearable devices (e.g., headphones / earbuds, smart glasses, HMDs, etc.) in a noisy environment may cause users to be unable to perceive the content with the correct loudness, timbre, spectral naturalness, externalization of spatial audio, and / or speech intelligibility by the content creator because of the auditory masking effect. Noise compensation aims to (partially) restore losses in loudness, naturalness, and timbre perception caused by environmental noise masking to give an equivalent perception of the content as in a quiet environment. Traditionally, noise compensation captures sounds via external microphones attached to devices as the environmental noise to be masked. But when a human wears the device to listen to a media (e.g., an audio), the wearer likely makes vocal sounds like voice, coughing, sneezing, humming songs, etc. Therefore, sound captured by external microphone(s) may be a mixture of environmental and vocal sounds. If the wearer’s vocal sound is mistakenly treated as the masking sound, noise compensation will lead to media fluctuation along with the wearer's self-voice, which is annoying for wearers. Additionally, if the environmental sound contains rich transient sound events with large dynamic ranges, a playback media fluctuation will occur and be perceived by wearers. In some other possible use cases / scenarios, context awareness (which may include environmental context) may be considered part of a personalization story that aims to leverage wearables coupled with portable electronic products to provide a tailored listening experience that contextually blends the real world with consumer listening experiences. However, regarding environmental context, important sound events in the ambient environment need to be considered on device embedded active noise cancellation (ANC) or headphones with aggressive physical isolation, especially for safety issues, which may include, for example, a baby crying in the living room, car-horn, siren alarms in the outdoor street, etc. In some further possible use cases / scenarios, people may tend to use wearable audio devices, including earbuds and headphones, in a lot of ambiences and scenes, so that they are isolated from the ambient sound. When the user wants to start a conversation with nearby people, they need to trigger pass-through mode and mute the playback media to hear the other people’s voices. Conversation awareness is to detect the time duration when the user has a conversation, so that the playback or noise cancellation can be adjusted to appropriate modes. However, during the conversation, the speech may usually be concatenated by the ambient sound, especially interference speech and intrusive noise, which may seriously affect the user’s conversation experience. Therefore, in view of all or at least some of the above illustrated potential problems / issues, there is a need for techniques (e.g., systems, methods, apparatuses, etc.) that are capable of improving the operation of such devices under real-world conditions, particularly with improved user experience. SUMMARY In view of the above, the present disclosure provides various systems, methods, apparatuses, and programs, as well as computer-readable storage mediums, having the features of the respective independent claims. The dependent claims relate to preferred embodiments. For example, the present disclosure provides a system, a method, an apparatus, and a program, as well as a computer-readable storage medium, for context-aware noise compensation. According to a first aspect of the present disclosure, there is provided a context-aware noise compensation system for a wearable device (or more generally, a mobile device) playing back media content (e.g., an audio, a video, or the like) to a user (e.g., the wearer of the wearable device) in a noisy (e.g., outdoor) environment. In particular, such wearable device (or mobile device in general) may comprise one or more sensors, which may include (but certainly not limited to) acoustic sensors (e.g., microphones, etc.), or non-acoustic sensors (e.g., accelerometers, etc.). Of course, as will be understood and appreciated by the skilled person, the wearable may comprise any other suitable sensor as well, depending on various implementations and / or circumstances. More particularly, the system may be configured to obtain a sound signal from the noisy environment. As will be discussed in more detail below, such sound signal may be obtained in any suitable manner, e.g., by making use of the available sensors of the wearable device. Then, the system may be configured to determine context information associated with the sound signal. Depending on various implementations and / or circumstances, the context information may comprise at least one of: vocal sound detection information or sound event detection information. In particular, the vocal sound detection information may be understood as being indicative of presence (or absence) of one or more vocal sounds of the user in the sound signal. On the other hand, the sound event detection information may be understood as being indicative of presence (or absence) of one or more environmental sound events in the sound signal. Finally, the system may be configured to optimize, based on the context information, the media content for playback by the wearable device. It may be worth mentioning that, although no specific module, entity, or component seems to have been explicitly mentioned above for the proposed system, as will be understood and appreciated by the skilled person, such system may be implemented to comprise respective suitable means configured for performing the respective features / steps as described throughout the present disclosure. Configured as proposed, broadly speaking, the present disclosure generally provides an efficient, flexible, yet reliable mechanism for enabling / supporting context-aware noise compensation for media playback in noisy environments, that is particularly capable of overcoming the fluctuation problem in conventional environmental masking techniques as mentioned above. In some example embodiments, the one or more vocal sounds of the user may comprise self- voice (or self-sound), such as speech, coughing, sneezing, or humming, of the user. Of course, as will be understood and appreciated by the skilled person, any other suitable self-voice or self-sound may be applicable as well, depending on various implementations and / or circumstances. This is not to be limited in the present disclosure. In some example embodiments, the one or more environmental sound events may comprise at least one transient sound event with a large dynamic range in the environment. For instance, the environmental sound events may be rich transient sound events with large (e.g., exceeding a predefined threshold / limit, or the like) dynamic ranges in environmental sound. In some possible examples, such environmental sound events may comprise at least one transient sound event that is at least 5 dB (or the like) louder than the noise floor. In some example embodiments, the wearable device may include: a headphone, an earbud, a smart glass, or a head-mounted device (HMD). Of course, as will be understood and appreciated by the skilled person, any other suitable wearable device, or, more generally, mobile device, may be considered applicable as well, depending on various implementations and / or circumstances. This is not to be limited in the present disclosure. In some example embodiments, the sound signal may be obtained by using one or more external microphones of the wearable device. As can be understood and appreciated by the skilled person, generally speaking, external microphones typically refer to those microphones facing outward and closest to the ear canal, which can be used to capture and estimate the (external) environmental sounds (e.g., environmental masking sources). In comparison, internal microphones typically refer to those attached in the ear and facing inward, which can be used to monitor the person’s listening state. While external microphones usually exist in many wearable devices, there are some wearable devices (e.g., open ear devices) that do not have in-ear (internal) microphones. In some example embodiments, the vocal sound detection information may be determined based on internal and external microphones of the wearable device jointly. By jointly leveraging sensor data obtained by both internal and external microphones, vocal sound detection can be achieved more efficiently and accurately. In some example embodiments, the one or more sensors may comprise a voice-pick-up (VPU) sensor (or simply referred to as VPU in short). As will be understood and appreciated by the skilled person, such VPU may be implemented in any suitable manner, e.g., in the form of a high speed accelerometer, or the like. This is not to be limited in the present disclosure. Accordingly, the vocal sound detection information may be determined based on the VPU sensor. Depending on various implementations and / or circumstances, such VPU sensor may be leveraged as an alternative to, or in combination with, the internal and / or external microphones as discussed above. In some example embodiments, the context information may further comprise information indicative of one or more sound sources separated by using the VPU sensor and the microphones. That is to say, in broad terms, the combination of VPU and microphones may be considered a powerful duo when it comes to sound source separation. The collaborative efforts of the combination of VPU and microphones can be leveraged to effectively separate and isolate different sound sources. In some example embodiments, the sound event detection information may be determined based on one or more external microphones of the wearable device. In some possible cases, the sound captured by the one or more external microphones may be used to implement sound event detection, and also noise floor detection. In some possible cases, if the wearable device comprises a plurality of external microphones, multiple external microphones may be used to implement for example beamforming to enhance sounds from specific directions. In some example embodiments, the optimization of the media content may involve: if one or more vocal sounds are detected or the detected one or more vocal sounds exceed a predetermined threshold level, estimating a current (e.g., of the current frame) level of the environment based on microphone data from the past (e.g., from the previous / latest (non- speech) frame); otherwise, determining the current level of the environment based on current microphone data. In some example embodiments, the optimization of the media content may involve: determining, based on the vocal sound detection information, a ratio of environmental sound sources to the sound signal. In some example embodiments, if the vocal sound detection information indicates absence of vocal sounds at a current frame, the ratio may be set to 1 for updating an excitation level of the environmental sound sources at the current frame; and if the vocal sound detection information indicates presence of one or more vocal sounds at the current frame, the ratio may be determined so as to keep the excitation level of the previous frame. In some example embodiments, the ratio ^ may be determined according to: ^1 ,^^ ^^ = 0^ (^, ^^) = ^^^(^^^,^^)^ , ^^ ^^ = 1, detection information, and ^^(^, ^^) denotes anexcitation level of the environmental sound sources at an ^-th frame at a center frequency ^^. Of course, as will be understood and appreciated by the skilled person, the ratio ^ may be determined in any other suitable manner, depending on various implementations and / or circumstances. For instance, in the above example embodiments, the determined ratio ^ may be seen as being frequency independent. In some other possible examples, the ratio ^ may be determined to be frequency (e.g., high-frequency) dependent as well. In some example embodiments, the ratio may be determined so as to mask a specific environmental sound source that has been separated. For instance, in some possible cases, the collaborative efforts of combining VPUs and microphones may be leveraged so as to effectively separate and isolate different sound sources, as illustrated above. In some example embodiments, the ratio ^ is determined according to:^(^, ^ ^ (^,^^) = ^ ^ ^)^^ , an excitation level of the environmental sound sources at an ^-thframe at a center frequency ^^, and ^^(^, ^^) represents the separated environmental soundsource. As such, noise compensation may be designed in such a manner to mask specific environmental sound sources only; for example, it can only mask the noise floor and ignore other sound events. Of course, as will be understood and appreciated by the skilled person, any other suitable masking may be possible as well, depending on various implementations and / or circumstances. In some example embodiments, the optimization of the media content may involve applying a time and / or frequency dependent gain to the media content. In some example embodiments, the gain may be determined such that, when being applied, a partial specific loudness of the media content in the noisy environment may be equal to a specific loudness for normal hearing and in quiet. In other words, broadly speaking, with the unified partial specific loudness of the target in one specific frequency band, the goal may be understood to optimize the gain applied to playback media that results in that the partial specific loudness of the target, given the noisy environment, should be equal to the specific loudness for normal hearing and in quiet. In some example embodiments, the gain ^ may be determined such that the following equation is fulfilled:^^(^^ ∙ ^ ^ $ ^ $ $! + # + ^ ∙ ^^) − (# + ^ ∙ ^^) & = ^^(^! + #) − #$&,where ^, #, and ' are constants, ^!denotes a running short-term estimate of excitation of the media content in quiet, ^ denotes a ratio of environmental sound sources to the sound signal, and ^^denotes an excitation level of the environmental sound sources. In some example embodiments, the gain ^ may be determined according to: 0 / ,^^ = ((^)*+),^+,*(+*-.∙^^), / ^+*-.∙^^^) . Of course, as will be understood and appreciated bythe skilled person, the gain ^ may be determined in any other suitable manner as well, depending on various implementations and / or circumstances. In some example embodiments, the determination of the context information may involve signal processing and / or artificial intelligence (AI) based techniques. As can be understood and appreciated by the skilled person, such AI based techniques may involve, among others, machine learning (ML) or deep learning (DL) -based techniques, or the like. According to a second aspect of the disclosure, there is provided a method of context-aware noise compensation for a wearable device (or more generally, a mobile device) playing back media content (e.g., an audio, a video, or the like) to a user (e.g., the wearer of the wearable device) in a noisy (e.g., outdoor) environment. As mentioned above, such wearable device (or mobile device in general) may comprise one or more sensors, which may include (but certainly not limited to) acoustic sensors (e.g., microphones, etc.), non-acoustic sensors (e.g., accelerometers, etc.), or any other suitable sensors, depending on various implementations and / or circumstances. In particular, the method may comprise obtaining a sound signal from the noisy environment (e.g., by making use of the available sensors of the wearable device, or the like). The method may further comprise determining context information associated with the sound signal, wherein the context information comprises at least one of: vocal sound detection information indicative of presence (or absence) of one or more vocal sounds of the user in the sound signal, or sound event detection information indicative of presence (or absence) of one or more environmental sound events in the sound signal. Finally, the method may comprise optimizing, based on the context information, the media content for playback by the wearable device. Configured as proposed, broadly speaking, the present disclosure generally provides an efficient, flexible, yet reliable mechanism for enabling / supporting context-aware noise compensation for media playback in noisy environments, that is particularly capable of overcoming the fluctuation problem in conventional environmental masking techniques as mentioned above. According to a third aspect of the present disclosure, there is provided an apparatus including a processor and a memory coupled to the processor. The processor may be adapted to cause the apparatus to carry out all steps according to any of the example methods described in the foregoing aspects. According to a fourth aspect of the present disclosure, there is provided a (computer) program. The (computer) program may include instructions that, when executed by a processor, cause the processor to carry out all steps of the example methods described throughout the present disclosure. According to a fifth aspect of the present disclosure, there is provided a computer-readable storage medium. The computer-readable storage medium may store the aforementioned (computer) program. The present disclosure also provides a system, a method, an apparatus, and a program, as well as a computer-readable storage medium, for selective sound event passthrough. According to a sixth aspect of the present disclosure, there is provided a system for selective sound event passthrough for a device (e.g., a wearable device, or more generally, a mobile device) playing back media content (e.g., an audio, a video, or the like) in an (e.g., outdoor) environment. As will be described in more detail below, as will also be understood and appreciated by the skilled person, such (wearable or mobile) device may comprise one or more sensors, which may include (but certainly not limited to) acoustic sensors (e.g., microphones, etc.), or non-acoustic sensors (e.g., accelerometers, etc.), or any other suitable sensors as well. More particularly, the system may comprise a sound event type determination module (or entity) configured for determining one or more target sound event types to be detected. The system may further comprise a sound event detection and segmentation module configured for detecting, in an environmental sound signal obtained from the environment, an environmental sound event matching any of the one or more target sound event types, and its respective onset and offset. As will be discussed in more detail below, such environmental sound signal may be obtained from the environment in any suitable manner, e.g., by making use of the available sensor(s) of the wearable device. Finally, the system may comprise a passthrough steering module configured for controlling remixing of the media content and the environmental sound event for playback via the device. Configured as proposed, broadly speaking, the present disclosure generally provides an efficient, flexible, yet reliable mechanism for enabling / supporting selective sound event passthrough to pass the specific sound events (e.g., those considered of importance to the user, especially for safety issues, such as, a baby crying in living room, car-horn, siren alarms in the outdoor street, or the like) together with playback media content, and more particularly, to contextually blend those real-world important sound events with playback media content by penetrating certain sound events, thereby greatly improving the overall experience during use of such (wearable or mobile) devices. In some example embodiments, the sound event type determination module may be configured for determining the one or more target sound event types to be detected based on user selection, pre- or default configuration, or automatically. For instance, in some possible examples, the sound event type determination module may be decided through the user’s personalized selection, for example, via a suitable user interface (UI) or the like. In some other possible examples, the specific sound event types may be determined based on the scene information, for example the target sound events may be car horn and siren on the street, or the like. In some further possible examples, the sound event types may be set by an application or by the mobile / wearable device (e.g., as a default setting, a preconfiguration, or the like). Moreover, in some possible cases, for example when (artificial intelligence) AI - based techniques are involved, the sound event types may be automatically determined (e.g., learned) during the course of classification or feature analysis. Incidentally, it may be worth mentioning that, even in the case of user selection for example via the UI, the list from which the sound event types are to be selected may not be always fixed, and may be extended on the way, for example by adding or recording (new / extra) sound event types that the user cares or is interested in now. Of course, as will be understood and appreciated by the skilled person, the sound event type may be determined or obtained by using any other suitable means as well, depending on various implementations and / or circumstances. In some example embodiments, the environmental sound signal may be obtained by using one or more microphones of the device. As will be described in detail below, the one or more microphones may comprise one or more external and / or internal microphones, depending on various implementations and / or circumstances. Broadly speaking, as will be understood and appreciated by the skilled person, external microphones typically refer to those microphones facing outward and closest to the ear canal, which can be used to capture and estimate the (external) environmental sounds (e.g., environmental masking sources). In comparison, internal microphones typically refer to those attached in the ear and facing inward, which can be used to monitor the person’s listening state. While external microphones usually exist in many wearable devices, there are some wearable devices (e.g., open ear devices) that do not have in-ear (internal) microphones. In some example embodiments, the detection of the environmental sound event and the respective onset and offset thereof may involve rule and / or AI -based techniques. As can be understood and appreciated by the skilled person, such AI based techniques may involve, among others, machine learning (ML) or deep learning (DL) -based techniques, or the like. In some example embodiments, the detection of the environmental sound event and the respective onset and offset thereof may involve feature extraction of the environmental sound signal. In some example embodiments, the feature extraction may involve: extracting full and / or subband features in a frequency domain, thereby capturing pitch and harmonic characteristic differences of the environmental sound signal; and / or extracting short and / or long -term temporal features in a time domain, thereby extracting time variation of the environmental sound signal. Of course, as will be understood and appreciated by the skilled person, any other suitable feature extraction may be exploited as well, depending on various implementations and / or circumstances. In some example embodiments, the passthrough steering module may be specifically configured to perform, for the controlling of the remixing of the media content and the environmental sound event: loudness analysis of the environmental sound event; and loudness analysis of the media content. In some example embodiments, the system may optionally further comprise a sound event enhancement module configured for enhancing the environmental sound event, for the remixing. In some example embodiments, the enhancement of the environmental sound event may involve digital signal processing (DSP) and / or artificial intelligence (AI) -based techniques. Of course, as will be understood and appreciated by the skilled person, any other suitable enhancement related technique may be exploited as well, depending on various implementations and / or circumstances. In some example embodiments, the enhancement of the environmental sound event may involve noise estimation and noise suppression. In some example embodiments, the noise estimation may involve: calculating a smoothing parameter based on the environmental sound event; calculating a smoothed power spectral of the environmental sound signal; searching for a minimum value of the smoothed power spectral in a time window; and calculating and updating the noise estimation based on the minimum value. Of course, as will be understood and appreciated by the skilled person, such noise estimation is merely provided as an illustrative example, and thus should not be understood to constitute a limitation of any kind. Depending on various implementations and / or circumstances, any other suitable technique / mechanism / algorithm may be applied as well. In some example embodiments, the device may be a mobile device, or a wearable device such as: a headphone, an earbud, a smart glass, or a head-mounted device (HMD). Of course, as will be understood and appreciated by the skilled person, any other suitable device may be considered applicable as well, depending on various implementations and / or circumstances. In some example embodiments, the one or more target sound event types to be detected may comprise at least one of: siren alarm, baby crying, car horn, door knocking, or house alarm. Of course, as will be understood and appreciated by the skilled person, any other suitable sound event type, which may include (but certainly not limited to) those of interest or importance to the user (e.g., safety related, or the like), may be set as target sound event type as well, depending on various implementations and / or circumstances. According to a seventh aspect of the disclosure, there is provided a method of selective sound event passthrough for a device (e.g., a wearable device, or more generally, a mobile device) playing back media content (e.g., an audio, a video, or the like) in an (e.g., outdoor) environment. In particular, the method may comprise determining one or more target sound event types to be detected. The method may further comprise detecting, in an environmental sound signal obtained (e.g., by making use of the available sensors of the wearable device, or the like) from the environment, an environmental sound event matching any of the one or more target sound event types, and its respective onset and offset. Finally, the method may comprise controlling remixing of the media content and the environmental sound event for playback via the device. Configured as proposed, broadly speaking, the present disclosure generally provides an efficient, flexible, yet reliable mechanism for enabling / supporting selective sound event passthrough to pass the specific sound events (e.g., those considered of importance to the user, especially for safety issues, such as, a baby crying in living room, car-horn, siren alarms in the outdoor street, or the like) together with playback media content, and more particularly, to blend those real-world important sound events with consumer listening experience, thereby greatly improving the overall experience during use of such (wearable or mobile) devices. In some example embodiments, the method may further comprise obtaining the environmental sound signal by using one or more microphones of the device. In some example embodiments, the method may further comprise, before the remixing, enhancing the environmental sound event. In some example embodiments, the method may involve one or more artificial intelligence (AI) based techniques. Such one or more AI based techniques may be applied to some or all of the (sub-)features / steps as described throughout the present disclosure. According to an eighth aspect of the present disclosure, there is provided an apparatus including a processor and a memory coupled to the processor. The processor may be adapted to cause the apparatus to carry out all steps according to any of the example methods described in the foregoing aspects. According to a ninth aspect of the present disclosure, there is provided a (computer) program. The (computer) program may include instructions that, when executed by a processor, cause the processor to carry out all steps of the example methods described throughout the present disclosure. According to a tenth aspect of the present disclosure, there is provided a computer-readable storage medium. The computer-readable storage medium may store the aforementioned (computer) program. The present disclosure further provides a system, a method, an apparatus, and a program, as well as a computer-readable storage medium, for supporting conversation awareness. According to an eleventh aspect of the present disclosure, there is provided a system for supporting conversation awareness during use of a device (e.g., a wearable device, or more generally, a mobile device) of a user. As will be described in more detail below, as will also be understood and appreciated by the skilled person, such (wearable or mobile) device may comprise one or more sensors, which may include (but certainly not limited to) acoustic sensors (e.g., microphones, etc.), or non-acoustic sensors (e.g., accelerometers, etc.), or any other suitable sensors as well. More particularly, the system may comprise a self-speech detection module (or entity) configured for detecting (or determining) a self-speech activity of the user (sometimes also simply referred to as self speech). The system may further comprise an external speech detection module configured for detecting (or determining) an external speech activity of one or more other participants of the conversation (sometimes also simply referred to as external speech). Finally, the system may comprise a conversation turn-taking tracking module configured for, based on the detection of the self-speech activity and / or the external speech activity, tracking an activity (or state) of the conversation. Configured as proposed, broadly speaking, the present disclosure generally provides an efficient, flexible, yet reliable mechanism for enabling / supporting intelligent conversation awareness, and potentially, also enhancement. For instance, as will be described in more detail below, this may involve, but certainly not limited to, being able to automatically detect the time duration of conversation based on the signal captured for example by microphone(s) and / or other suitable sensor(s) of the device, and optionally, to enhance the target speech and remove annoying intrusive noises during conversation mode. In some example embodiments, the self-speech activity of the user may be detected based on at least one of: a microphone signal of the device, a sensor signal of the device, or a media signal played back by the device. As will be described in detail below, the microphone may be an external microphone or an internal microphone, depending on various implementations and / or circumstances. Broadly speaking, as will be understood and appreciated by the skilled person, external microphones typically refer to those microphones facing outward and closest to the ear canal, which can be used to capture and estimate the (external) environmental sounds (e.g., environmental masking sources). In comparison, internal microphones typically refer to those attached in the ear and facing inward, which can be used to monitor the person’s listening state. While external microphones usually exist in many wearable devices, there are some wearable devices (e.g., open ear devices) that do not have in-ear (internal) microphones. Of course, as will be understood and appreciated by the skilled person, any other suitable input may be fed to the self speech detection module as well, in order to facilitate the detection of the self speech activity, depending on various implementations and / or circumstances. In some example embodiments, the device may comprise an internal microphone capable of capturing self-speech of the user, thereby obtaining an internal microphone signal. The device may also comprise an external microphone capable of capturing self-speech of the user, thereby obtaining an external microphone signal that further comprises ambient sound. More particularly, the self-speech activity of the user may be detected (or determined) based on the internal microphone signal and the external microphone signal. In some example embodiments, the internal microphone signal may further comprise a media signal played back by the device. Particularly, acoustic echo cancellation (AEC), or similar, may be applied to the internal microphone signal for suppressing or attenuating the media signal, before the internal microphone signal is used for the detection of the self-speech activity of the user. In some example embodiments, the detection of the self-speech activity of the user may involve digital signal processing (DSP) and / or artificial intelligence (AI) -based techniques. Of course, as will be understood and appreciated by the skilled person, any other suitable technique may be exploited as well, depending on various implementations and / or circumstances. In some example embodiments, the device may comprise two or more external microphones on one side of the user. In other words, the two or more external microphones may be placed (arranged) on the same side (e.g., left or right) relative to (or in reference to) the user. In particular, each of the two or more external microphones may be configured for obtaining a respective external microphone signal that comprises self-speech of the user and ambient sound (of the surrounding / environment). More particularly, the detection of the self-speech activity of the user may involve one or more spatial selection-based techniques, such as beamforming (or the like). In some example embodiments, the device may comprise a respective external microphone on each side relative to the user, each configured for obtaining a respective external microphone signal that comprises self-speech of the user and ambient sound. In other words, in such cases, a separate external microphone may be placed (arranged) on each side (e.g., one external microphone on the left and another one on the right) relative to (or in reference to) the user. In particular, the self-speech activity of the user may be detected based on binaural speech information derived from the external microphone signals, or the like. In some example embodiments, the binaural speech information may comprise: interaural phase difference (IPD), interaural time difference (ITD), and / or interaural level difference (ILD). Of course, as will be understood and appreciated by the skilled person, any other suitable metric or measure may be exploited as well, depending on various implementations and / or circumstances. In some example embodiments, the device may comprise (only) one microphone. In particular, in such cases, if the microphone is an external microphone configured for obtaining an external microphone signal that comprises self-speech of the user and ambient sound, the detection of the self-speech activity of the user may involve voice identification and / or voice authentication. On the other hand, if the microphone is an internal microphone configured for obtaining an internal microphone signal that comprises self-speech of the user and a media signal played back by the device, the detection of the self-speech activity of the user may involve acoustic echo cancellation (AEC) for suppressing or attenuating the media signal. In some example embodiments, the device comprises a voice pick-up (VPU) sensor. The VPU sensor may for example be bone conduction based, high speed accelerometer based, or the like. In particular, the self-speech activity of the user may be detected based on data obtained by the VPU sensor. In some example embodiments, the external speech activity may be detected based on signals obtained from one or more microphones (e.g., external microphones) of the device. In some example embodiments, the detection of the external speech activity may involve extracting at least one of: time, frequency, or spatial, features from the signals. In some example embodiments, the device may comprise a microphone array. In such cases, the external speech activity may be detected based on a configuration or layout of the microphone array. In some example embodiments, if the microphone array comprises two or more microphones on one (the same) side of the user, the detection of the external speech activity may involve beamforming or similar techniques. On the other hand, if the microphone array comprises a respective microphone on each side relative to the user, the detection of the external speech activity may involve binaural auditory based noise suppression for enhancing sound from a target direction but suppressing sounds from other directions. In some example embodiments, a speech detection (sometimes also referred to as speech / non- speech detection) may be performed on a single-channel signal obtained by one microphone of the microphone array, before triggering the beamforming. In such cases, the proposed system may be seen as structured in a multi-stage framework to detect speech from the target direction, thereby improving the computational efficiency of external speech detection. In some example embodiments, the device may comprise one (or even more) camera. In such cases, the external speech activity may be detected further based on face and / or lip movement detection by using the camera. Of course, as will be understood and appreciated by the skilled person, any other suitable camera based technique may be exploited as well for facilitating the detection of the external speech activity, depending on various implementations and / or circumstances. In some example embodiments, the tracking of the activity of the conversation may involve determining at least one of: onset, offset, or states of the conversation. In some example embodiments, the states of the conversation may comprise: a self-speech state, an external speech state, a double talk state, and a mutual silence state. In some example embodiments, transition between the states of the conversation may be based on the detection of the self-speech activity and / or the external speech activity. In some example embodiments, the transition may be further based on a statistic or heuristic rule -based probability. In some example embodiments, the mutual silence state may be configured with a hold on timer that is heuristically or adaptively configurable. For instance, depending on various implementations, the maximum hold on (waiting) time may be heuristically set or adaptively updated based on the history of speech burst lengths. In some possible implementations, if the user keeps speaking long sentences, then the hold on time can be set a bit longer; whilst the short and quick conversation should have a relatively shorter hold on time. In some possible examples, during the hold on time, the conversation existence probability may also be adjusted, for example according to signal(s) from microphones, cameras, sensors, or the like, so that the system can control the playback media playback and pass-through level to improve the user’s listening experience and shorten the waiting time. In some example embodiments, the system may optionally further comprise a conversation enhancement module configured for enhancing the conversation. In some example embodiments, the enhancement of the conversation may involve at least one of: enhancing speech during the conversation, suppressing ambience or background noise, attenuating interference speech, or stop / replay of media content played back by the device. Of course, as will be understood and appreciated by the skilled person, any other suitable enhancement related processing / technique may be applied as well, depending on various implementations and / or circumstances. In some example embodiments, the enhancement of the conversation may involve remixing the conversation with virtual media based on deep learning based speech isolation. In some example embodiments, the enhancement of the conversation may involve digital signal processing (DSP) and / or artificial intelligence (AI) -based techniques. Of course, as will be understood and appreciated by the skilled person, any other suitable enhancement related technique may be exploited as well, depending on various implementations and / or circumstances. In some example embodiments, the device may be a mobile device, or a wearable device such as: a headphone, an earbud, a smart glass, or a head-mounted device, HMD. Of course, as will be understood and appreciated by the skilled person, any other suitable device may be considered applicable as well, depending on various implementations and / or circumstances. According to a twelfth aspect of the disclosure, there is provided a method of supporting conversation awareness during use of a device (e.g., a wearable device, or more generally, a mobile device) of a user. In particular, the method may comprise detecting (or determining) a self-speech activity of the user (sometimes also simply referred to as self speech). The method may further comprise detecting (or determining) an external speech activity of one or more other participants of the conversation (sometimes also simply referred to as external speech). Finally, the method may comprise tracking an activity (or state) of the conversation, based on the detection of the self- speech activity and / or the external speech activity. Configured as proposed, broadly speaking, the present disclosure generally provides an efficient, flexible, yet reliable mechanism for enabling / supporting intelligent conversation awareness, and potentially, also enhancement. For instance, as will be described in more detail below, this may involve, but certainly not limited to, being able to automatically detect the time duration of conversation based on the signal captured for example by microphone(s) and / or other suitable sensor(s) of the device, and optionally, to enhance the target speech and remove annoying intrusive noises during conversation mode. In some example embodiments, the method may further comprise enhancing the conversation. In some example embodiments, the method involves one or more artificial intelligence (AI) - based techniques. Such one or more AI based techniques may be applied to some or all of the (sub-)features / steps as described throughout the present disclosure. According to a thirteenth aspect of the present disclosure, there is provided an apparatus including a processor and a memory coupled to the processor. The processor may be adapted to cause the apparatus to carry out all steps according to any of the example methods described in the foregoing aspects. According to a fourteenth aspect of the present disclosure, there is provided a (computer) program. The (computer) program may include instructions that, when executed by a processor, cause the processor to carry out all steps of the example methods described throughout the present disclosure. According to a fifteenth aspect of the present disclosure, there is provided a computer- readable storage medium. The computer-readable storage medium may store the aforementioned (computer) program. It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus (or system), and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) are understood to likewise apply to the corresponding apparatus (or system), and vice versa. BRIEF DESCRIPTION OF DRAWINGS Example embodiments of the disclosure are explained below with reference to the accompanying drawings, wherein Fig.1 schematically illustrates an example block diagram of media playback in a noisy environment according to some example embodiments of the present disclosure, Fig.2 schematically illustrates an example of a method of context-aware noise compensation according to some example embodiments of the present disclosure, Fig.3 schematically illustrates another example of a method of context-aware noise compensation according to some example embodiments of the present disclosure, Fig.4 schematically illustrates an example of a typical use case of a wearable device by a user, Fig.5 schematically illustrates an example block diagram of selective sound event passthrough according to some example embodiments of the present disclosure, Fig.6 schematically illustrates examples of environment signal sources according to some example embodiments of the present disclosure, Fig.7 schematically illustrates an example of personalized selection of a user according to some example embodiments of the present disclosure, Fig.8 schematically illustrates examples of sound event detection and segmentation methods according to some example embodiments of the present disclosure, Fig.9 schematically illustrates an example of feature extraction according to some example embodiments of the present disclosure, Fig.10 schematically illustrates an example of a machine learning based selective sound event detection model training according to some example embodiments of the present disclosure, Fig.11 schematically illustrates an example of a machine learning based selective sound event detection model inferencing according to some example embodiments of the present disclosure, Fig.12 schematically illustrates an example of a deep learning based selective sound event detection model training according to some example embodiments of the present disclosure, Fig.13 schematically illustrates an example of a deep learning based selective sound event detection model inferencing according to some example embodiments of the present disclosure, Fig.14 schematically illustrates an example of sound event enhancement according to some example embodiments of the present disclosure, Fig.15 schematically illustrates an example of passthrough steering according to some example embodiments of the present disclosure, Fig.16 schematically illustrates an example of a method of selective sound event passthrough according to some example embodiments of the present disclosure, Fig.17 schematically illustrates another example of a method of selective sound event passthrough according to some example embodiments of the present disclosure, Fig.18 schematically illustrates an example of a microphone placement of a true wireless stereo (TWS) earbud, Fig.19 schematically illustrates an example block diagram of a conversation awareness system according to some example embodiments of the present disclosure, Fig.20 schematically illustrates an example of self speech detection according to some example embodiments of the present disclosure, Fig.21 schematically illustrates another example of self speech detection according to some example embodiments of the present disclosure, Fig.22 schematically illustrates a further example of self speech detection according to some example embodiments of the present disclosure, Fig.23 schematically illustrates an example of external speech detection according to some example embodiments of the present disclosure, Fig.24 schematically illustrates another example of external speech detection according to some example embodiments of the present disclosure, Fig.25 schematically illustrates an example of conversation turn-taking tracking according to some example embodiments of the present disclosure, Fig.26 schematically illustrates an example of a method of conversation awareness according to some example embodiments of the present disclosure, Fig.27 schematically illustrates another example of a method of conversation awareness according to some example embodiments of the present disclosure, and Fig.28 is a schematic block diagram of an example electronic device or architecture. DETAILED DESCRIPTION The Figures (Figs.) and the following description relate to preferred embodiments by way of illustration only. It should be noted that from the following discussion, alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be employed without departing from the principles of what is claimed. Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying figures. It is noted that wherever practicable, similar or like reference numbers may be used in the figures and may, unless indicated otherwise, indicate similar or like functionality, such that repeated description thereof may be omitted for reasons of conciseness. The figures depict embodiments of the disclosed system (or method) for purposes of illustration only. One skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein. It is also to be noted that while some of the examples shown in the figures and described below may seem to explicitly make reference to particularly wearable devices (e.g., headphones, earbuds, etc.), they are merely provided as possible examples for easy illustration purposes only, and thus should not be understood to constitute limitations of any kind. As will be understood and appreciated by the skilled person, any other suitable type of wearable device, or, more generally, mobile device, may be considered applicable as well, depending on various implementations and / or circumstances. As briefly mentioned above, nowadays, wearable devices (or, more generally, mobile devices), such as headphones, earbuds, smart glasses, and head-mounted devices (HMDs), etc., are increasingly used in noisy and variable environments, for example for media (e.g., audio, video) playback. In a broad sense, users / customers (e.g., wearers of such mobile / wearable devices) may typically expect consistent media playback, awareness of relevant ambient sounds (or sound events), natural conversational support, or the like, which may however pose various challenges or issues in practice. Generally speaking, the present disclosure seeks to propose techniques (e.g., systems, methods, apparatuses, etc.) to address at least some or all of those potential challenges / issues, thereby improving the operation of such (wearable / mobile) devices under real-world conditions, particularly with improved user experience. Before going into detail, it may be worthwhile to mention that the following description of the present disclosure may appear to be structured in such a manner that several challenges / issues seem to have been discussed and addressed individually. However, such structure is merely for the sake of illustration, and thus should not be considered to constitute a limitation of any kind. As will be understood and appreciated by the skilled person, the respective techniques proposed therefor throughout the present disclosure may certainly be suitably applied in combination as well, depending on various implementations and / or circumstances. Context-Aware Noise Compensation As mentioned above, media playback by various wearable devices, such as headphones / earbuds, smart glasses, HMD, etc., in a noisy environment, may cause users to be unable to perceive the content with the correct loudness, timbre, spectral naturalness, externalization of spatial audio, and speech intelligibility by the content creator because of the auditory masking effect. Broadly speaking, noise compensation based techniques aim to (at least partially) restore losses in loudness, naturalness, and timbre perception caused by environmental noise masking to give an equivalent perception of the content as in a quiet environment. Traditionally, noise compensation may be understood to capture sounds via external microphones attached to devices as the environmental noise to be masked. However, in some possible cases, when a human wears the device to listen to media (e.g., an audio), the wearer (or user) may likely make vocal sounds like voice, coughing, sneezing, humming songs, etc. Therefore, sounds captured by external microphone(s) may be a mixture of environmental and vocal sounds. If the wearer's vocal sound is mistakenly treated as the masking sound, noise compensation will lead to media fluctuation along with the wearer’s self-voice (or self-sound), which would be annoying for the wearer. Additionally, in some possible cases, if the environmental sound contains rich transient sound events with large dynamic ranges (5 dB louder than the noise floor, or the like), a playback media fluctuation may occur and be perceived by the wearer. In view of the above, in broad terms, particularly in an attempt to overcome the fluctuation problem, the present disclosure may be understood to first propose a context-aware noise compensation method for listening and optimizing media in noisy environments, and a system to optimize playback audio for environmental masking. Loudness Model in Quiet As will be understood and appreciated by the skilled person, the perceived loudness may be described by means of excitation and specific loudness models. Such models typically start with the calculation of an excitation level in a set of frequency bands. Typically, the excitation level may be understood to vary across frequency and time for time-varying sounds. For simplicity and without loss of generality, the process for a single frequency band will be described below as an illustrative example. The instantaneous excitation level ^!2for a targetinput signal at the ^th time frame 3(^, ^) represented in the frequency domain may be givenby: ^;!2(^, ^^) = 42 ‖^6(^)^789(^, ^^)3(^, ^)‖^:^ , (1)with ^6(^) denoting auditory filter transfer function centered at frequency ^^, and ^ representing the frequency in Hz. In the exemplary loudness model, it may be assumed that a running average of the instantaneous excitation is taken to give the loudness of a sound. The averaging process in the model may be understood to resemble the operation of an automatic gain control system with separate attack and release times (<=and <8). The loudness estimate may increase rapidly when the level of a sound suddenly increases, but the estimate may decrease more slowly when the sound level decreases, which may be understood to reflect the fact that the loudness of a sound increases rapidly when the sound is first turned on, but the loudness impression decays more slowly when the sound is turned off.Assuming that ^!(^, ^^) is defined as the running short-term estimate of excitation at the timecorresponding to the ^th time frame (updated every <>). If ^!2(^, ^^) > ^!(^, ^^), then^!(^, ^^) = '= ∙ ^!2(^, ^^) + (1 − '=) ∙ ^!(^ − 1, ^^), (2)where '=is a (e.g., ABC'= = @BD . (3)If ^!2(^, ^^) ≤ ^!(^, ^^), then^!(^, ^^) = '8 ∙ ^!2(^, ^^) + (1 − '8) ∙ ^!(^ − 1, ^^), (4)where '8is a (e.g., ABC'8 = @ BF . (5)The specific loudness GH! (e.g., the loudness associated with signal 3(^, ^) in the frequencyband of interest at ^th frame) may then be computed according to: GH $!(^, ^^) = ^(^!(^, ^^) + #) − ^#$, (6)with constants ^, #, and '. A typical value of ' equals 0.3 (or may be set to any other suitable number). Constants ^, # may be used to calibrate the model. For instance, in some possible examples, one common calibration rule may be to set the constant ^ such that the sum of specific loudness values G!Hacross frequency bands ^^amounts to 1 sone for a 1-kHz sinusoidal input signal at a level of 40 dB sound pressure level (SPL). Loudness Model in Environment Next, loudness model in a (noisy) environment will be discussed. In particular, as will be understood and appreciated by the skilled person, masking signals (e.g., environmental sound(s)) can reduce the perceived loudness of a target signal (e.g., the media content being played back). One example is the masking effect of environmental sound on the perception of a media signal (target). In the following, it is assumed that the masking signal in isolation may cause an excitation level ^^. The loudness of a target in the presence of a masker may be denoted as the partial specific loudness G′!,^. The partial specific loudness of the target signal in the presence of the masker may then be given by the difference between the specific loudness of target and masker, minus the specific loudness of the masker in isolation according to: G′ $!,^ = ^^(^! + # + ^^) − (# + ^^)$&, (7)where ^ and ^^have With the above loudness model being described, reference is now made to Fig.1, which schematically illustrates an example block diagram 1000 of media playback in a noisy environment according to some example embodiments of the present disclosure. In particular, as the use case for media playback being visualized in Fig.1, media content 1310 (e.g., audio, video, or the like) from the playback device 1300 (e.g., PCs, smartphones, wearable devices, or the like) may be transmitted to the listener’s wearable / mobile device(s) 1100, like headphones / earbuds, smart glasses, head-mounted devices or the like, for example via a wired or wireless transmission. Depending on various implementations of the wearable / mobile device, there may be multiple types of sensors attached to the wearable / mobile devices, which may include (but certainly not limited to) acoustic based sensors (e.g., microphones, etc.), or non-acoustic based sensors (e.g., accelerometers, etc.). For example, such sensors may include a voice-pick-up (VPU) sensor 1120 that may for example rely on bone conduction (or any other suitable means) to accurately and reliably read a person’s skull vibrations and translate them into sound. The sensors may also include one or more external microphones 1110 facing outward and closest to the ear canal that can be used to capture and estimate the environmental masking sources 1220. Similarly, the sensors may include one or more internal microphones 1110 attached in the ear and facing inward that can be used to monitor the person's listening state. While external microphones usually exist in many wearable devices, in some possible cases, there may be some wearable devices (e.g., open ear devices, or the like) that do not include in-ear (internal) microphones. As will be understood and appreciated by the skilled person, such sensors may be used individually or jointly, depending on various implementations and / or circumstances. For instance, the VPU may be used to implement vocal sound detection. The vocal can be the voice (e.g., speech) or other sounds like coughing, sneezing, humming songs, etc., of the wearer / user. In some possible cases, for devices (e.g., some headphones) without VPU attached, internal and external microphones may be jointly leveraged to implement vocal sound detection. Further, external / environment sound captured by external microphone(s) may be used to implement sound event detection and / or noise floor detection. In some possible cases, multiple external microphones may be utilized to implement beamforming to enhance sounds from specific directions. In some possible examples, the combination of VPU and microphones may be considered a powerful duo when it comes to sound source separation. In particular, the collaborative efforts of the VPU and microphones used in combination may effectively separate and isolate different sound sources. As such, various context-aware information 1210 may be generated based on different sensor configurations (combinations). Broadly speaking, the context-aware information 1210 may be seen as extra information for noise compensation 1230 that modifies, e.g., by applying a gain 1240 or the like, (the audio of) the media content signal 1310 in response to the media signal 1310 and the environmental masking sources 1220. This optimized audio 1320 is finally played back by wearable devices 1100. As will become apparent with the description below, depending on various implementations and / or circumstances, the context information may comprise at least one of: vocal sound detection information or sound event detection information. In particular, the vocal sound detection information may be generally understood as being indicative of presence (or absence) of one or more vocal sounds of the user in the sound signal captured by the microphone(s). On the other hand, the sound event detection information may be generally understood as being indicative of presence (or absence) of one or more environmental sound events (e.g., rich transient sound events with large dynamic ranges in the environment) in the sound signal. In some possible examples, in order to compensate for the environmental noise, a time and / orfrequency-dependent gain ^(^, ^^) 1240 may be applied to the playback media 1310 to maskthe environmental noise. Accordingly, the resulting partial specific loudness of the target G′!,^in one specific frequency band may then be given by: G′! (^, ^ ) = ^^(^^,^ ^ (^, ^^) ∙ ^!(^, ^^) + # + ^^(^, ^^) ∙^ (8)^ ^where ^ events, to the whole external microphone captured sound. When a human wears a wearable device to listen to media, the wearer / user may likely make vocal sounds, such as voice, speech, coughing, sneezing, humming songs, or the like. Therefore, the external microphone(s) may be configured to capture sound that is a mixture of vocal and environmental sounds of the wearer / user. Suppose the noise compensation tries tomask the whole captured sound (^ = 1). In that case, there would be annoying medialoudness fluctuation because the playback media may be modulated by the vocal sound of the wearer / user, which is not the environmental sound. With the vocal sound detection information of the wearer / user, this issue may be solved, and ^ could then be given according to: ^1 ,^^ ^^ = 0^ ^^^ (9)where ^^may be That is, if ^^ = 0, it may be understood to correspond to the whole external microphonecapture sound at the current frame as noise to be masked. The excitation at this time framewould then be updated based on ^ = 1. On the other hand, if ^^ = 1, it may be understood tocorrespond to the external microphone captured sound at the current frame is a mixture of vocal sound and environmental sound. Then the excitation would keep the value of the previous (non-speech) frame to avoid fluctuation caused by tracking the vocal sound of the wearer / user. It may be worth mentioning that, as will be understood and appreciated by the skilled person, the ratio ^ may be determined in any other suitable manner, depending on various implementations and / or circumstances. For instance, in the above examples, the determined ratio ^ may be seen as being frequency independent. In some other possible examples, the ratio ^ may be determined to be frequency (e.g., high-frequency) dependent as well. Furthermore, the context-aware information 1210 may comprise context-aware sound event detection ^7that can for example indicate rich transient sound events with large dynamic ranges in the environment (e.g., 5 dB louder than the noise floor), or the like. Likewise, in some possible examples, this can also be used as ^^to mitigate playback media fluctuation caused by environmental sound events. As indicated above, the potential collaborative efforts of combining VPUs and microphones may be leveraged to effectively separate and isolate different sound sources. In such cases, noise compensation can be designed to mask specific environmental sound sources only; for example, it can only mask the noise floor and ignore other sound events. Accordingly, ^ may then be given by: ^(^, ^ ^ (^,^^) = ^ ^ ^)^^ , (10)where ^^may be understood to floor. With the unified partial specific loudness of the target G′!,^in one specific frequency band,the goal may be understood to optimize the gain ^(^, ^^) applied to playback media thatresults in G′!,^=G!H. For instance, in some possible example, this may mean that the partial specific loudness of the target, given the noisy environment, should be equal to the specific loudness for normal hearing and in quiet: ^^(^^ ∙ ^ ^ $ ^ $ $! + # + ^ ∙ ^^) − (# + ^ ∙ ^^) & = ^^(^! + #) − #$&, (11)For simplicity and without loss of generality, ^ and ^^may be omitted, thereby the optimized ^ can be derived by:^^ ^(^ + #)$ − #$! + (# + ^^ ∙ ^^)$&^ / $ − # + ^^ ∙ ^= ^^ . (12)!It is to be thus, should a any and appreciated by the skilled person, the gain ^ may be determined in any other suitable manner, depending on various implementations and / or circumstances. Fig.2 is a flowchart of an example method 2000 for context-aware noise compensation for a wearable device. The method 2000 may be performed by an electronic processor (for example, of the electronic apparatus 2800 of Fig.28) or any other suitable apparatus, which may be configured to perform the method via execution of machine-executable instructions. The various process blocks illustrated in Fig.2 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. In particular, at step S2001, the method 2000 may detect one or more sounds in an external environment in reference to the wearable device. For example, the electronic processor 2801 may detect, via a microphone, one or more sounds in an external environment in reference to the wearable device. At step S2002, the method 2000 may extract contextual (or context) information corresponding to the one or more sounds. For example, the electronic processor 2801 may extract contextual information corresponding to the one or more sounds (e.g., if the one or more sounds contain vocal sounds and / or environment noises). At step S2003, the method 2000 may, based on the contextual information, perform noise compensation, via the wearable device, to mask at least one of the one or more sounds. For example, the electronic processor 2801 may perform noise compensation, by applying a gain to the media being played back via the wearable device, to mask a sound (e.g., environmental noise). Fig.3 schematically illustrates another example of a method 3000 of context-aware noise compensation according to some example embodiments of the present disclosure. In particular, the method may be for context-aware noise compensation for a wearable device (or more generally, a mobile device) playing back media content (e.g., an audio, a video, or the like) to a user (e.g., the wearer of the wearable device) in a noisy (e.g., outdoor) environment. More particularly, the method 3000 may comprise, at step S3001, obtaining a sound signal from the noisy environment (e.g., by making use of the available sensors of the wearable device, or the like). The method 3000 may further comprise, at step S3002, determining context information associated with the sound signal, wherein the context information comprises at least one of: vocal sound detection information indicative of presence of one or more vocal sounds of the user in the sound signal, or sound event detection information indicative of presence of one or more environmental sound events in the sound signal. Finally, the method 3000 may comprise, at step S3003, optimizing, based on the context information, the media content for playback by the wearable device. To summarize the above, broadly speaking, the proposed techniques generally seek to provide an efficient, flexible, yet reliable mechanism for enabling / supporting context-aware noise compensation for media playback in noisy environments, that is particularly capable of overcoming the fluctuation problem in conventional environmental masking techniques as mentioned above. Accordingly, in some examples, the present disclosure provides a context-aware noise compensation framework used for headphones / earbuds, smart glasses, or head-mounted devices, to solve the playback media fluctuation problem caused by wearers’ vocal sounds, such as voice, coughing, sneezing, humming songs, etc., or rich transient sound events with large dynamic ranges in environmental sound, comprising: • Environmental sound perception via external microphones attached to headphones / earbuds, smart glasses, or head-mounted devices. • Context-aware information extraction contains wearer’s vocal sound detection, environmental sound event detection, or environmental sound source separation. • Context-aware noise compensation optimizes media in noisy environments to mask environmental sources to restore losses in loudness, naturalness, and timbre perception caused by environmental source masking. In some examples, the external microphones capturing environment sound should be facing outward and closest to the ear canal to approximate human hearing. In some examples, context-aware vocal sound detection can be implemented based on the Voice-Pick-Up (VPU) sensor data. For devices (some headphones, for example) without VPU attached, we can jointly leverage internal and external microphones to implement vocal sound detection. In some examples, context-aware sound event detection can be implemented based on external microphone captured sound. Additionally, multiple external microphones can implement beamforming to enhance sounds from specific directions. In some examples, context-aware environmental sound source separation can be implemented using a combination of VPUs and microphones. In some examples, a context-aware noise compensation system is proposed to apply a time and frequency-dependent gain ^ to the playback media to mask the environmental source determined by the context-aware information and solve the playback media fluctuation problem. In some examples, a time and frequency-dependent gain ^ is modeled to apply to the external microphone captured sound to represent the ratio of the environmental source to the whole captured sound.In some examples, context-aware vocal sound detection ^^ is binary value. If ^^ = 0, itcorresponds to the whole external microphone capture sound at current frame as noise to bemasked. The excitation at this time frame will be updated based on ^ = 1. If ^^ = 1, itcorresponds to the external microphone captured sound at current frame being a mixture of vocal sound and environmental sound, then the excitation will keep the value of the pervious frame to avoid fluctuation caused by tracking the wearer’s vocal sound. In some examples, context-aware sound event detection ^7can indicate rich transient sound events with large dynamic ranges in the environment. Likewise, it can be used as ^^to mitigate playback media fluctuation caused by environmental sound events. In some examples, context-aware environmental sound source separation can separate environmental sound sources. In this condition, ^ can be directly calculated by comparing the environmental sound source to the whole external microphone capture sound. An environmental sound source can be a noise floor, and then playback media can only mask the noise floor and ignore other sound events. Therefore, media fluctuation caused by rich transient sound events with large dynamic ranges in the environment can be mitigated.In some examples, time and frequency-dependent optimized gain ^(^, ^^) can be derivedresulting in G′!,^=G!H, e.g., the partial specific loudness of the media given the noisy environment being equal to the loudness of media in quiet. In some examples, the vocal sound detection includes self-voice, coughing, sneezing, humming songs, etc. In some examples, the techniques described herein relate to a context-aware noise compensation method for a wearable device, the method including: detecting one or more sounds in an external environment in reference to the wearable device; extracting contextual information corresponding to the one or more sounds; and based on the contextual information, performing noise compensation, via the wearable device, to mask at least a sound of the one or more sounds. In some examples, the techniques described herein relate to a method, wherein the detecting includes detecting, via a microphone, the one or more sounds in the external environment in reference to the wearable device. In some examples, the techniques described herein relate to a method, wherein the wearable device includes the microphone. In some examples, the techniques described herein relate to a method, wherein the microphone is configured to approximate human hearing. In some examples, the techniques described herein relate to a method, wherein the microphone is configured to face outward and in proximity to an ear canal of a user of the wearable device, in reference to when the user wearing the wearable device. In some examples, the techniques described herein relate to a method, wherein the wearable device includes at least one of headphones, earbuds, smart glasses, and / or head-mounted devices. In some examples, the techniques described herein relate to a method, wherein the contextual information includes at least one of vocal sound detection information, environmental sound event detection information, and / or environmental source separation information. In some examples, the techniques described herein relate to a method, wherein the one or more sounds include a mixture of vocal sounds of a user of the wearable device and external environment sounds. In some examples, the techniques described herein relate to a method, when the one or more sounds correspond to a vocal sound, the detecting includes detecting the one or more sounds based on Voice-Pick-Up sensor data or based on an internal microphone and an external microphone. In some examples, the techniques described herein relate to a method, wherein the wearable device includes a plurality of microphones, and wherein the detecting includes beamforming the microphones to determine sounds from specific areas of the external environment. In some examples, the techniques described herein relate to a method, wherein performing noise compensation, via the wearable device, to mask at least a sound of the one or more sounds includes: applying a gain to playback media of the wearable device to mask the at least a sound of the one or more sounds. In some examples, the techniques described herein relate to a method, wherein the gain includes a time and frequency-dependent gain. Selective Sound Event Passthrough As mentioned above, true wireless stereo (TWS) earbuds, headphones, smart glasses, etc., may be foreseen to be the most important methods to produce a great audio experience for portable electronic products, such as mobile or wearable devices. Because of the immersive media content, embedded active noise cancellation (ANC) or headphones with aggressive physical isolation, the user may be totally immersed in the world of the playback content. Accordingly, the present disclosure also seeks to solve the problem of blending the real-world important sound events with consumer listening experiences. In broad terms, the present disclosure attempts to develop a selective sound event passthrough algorithm to pass the specific sound events together with playback media content to improve the whole hearing experience. For instance, such selective sound events may be those considered important to users, especially for safety issues, such as a baby crying in the living room, a car-horn, siren alarms in the outdoor street, or the like. As also discussed above, generally speaking, context awareness may be seen as part of the personalization story that aims to leverage wearables coupled with portable electronic products to provide a tailored listening experience that contextually blends the real world with consumer listening experiences. Context awareness includes environmental context. Regarding environmental context, important sound events in the ambient environment need to be considered on device embedded ANC or headphones with aggressive physical isolation, for example, sirens / car horns / baby crying, etc. Fig.4 schematically shows a typical use case where users are immersed in the audio bubble and ignore important external sound events. According to the techniques proposed in the present disclosure, based on the user’s personalized selection, passing through the specific sound events together with playback media content will significantly improve the whole hearing experience, especially for safety issues, such as car-horn, siren alarms, or the like. In broad terms, as the present disclosure aims to provide a tailored listening experience that contextually blends the real world with playback media content by penetrating certain sound events, it would be expected to deal with at least the following aspects: • According to the user’s personalized selection, the proposed algorithm can timely and accurately locate the sound events that users care about. • The proposed algorithm can pass through clean sound events and guarantee that the quality of blended playback content is high. • The target sound event types are easily expanded. Accordingly, in the present disclosure, generally speaking, it is proposed to collect environmental signals for problem analysis and to develop a selective sound event pass- through algorithm that can penetrate certain sounds for a better hearing through experience. Thereby, the present disclosure provides, among others, a method and a system for automatic detection, segmentation, enhancement, and passthrough of selective sound event(s) in a real environment for portable electronic products to provide a tailored listening experience that contextually blends the real world with consumer media content. As will be discussed below in more detail, in a broad sense, the proposed system architecture may be understood to comprise one or more of the following: • sound event type determination module, which can define the target sound event according to the user’s personalized selection or adaptively; • selective sound event detection and segmentation module, which can figure out the sound event type and its onset and offset; • sound event enhancement module, which can remove background noise and get a clean sound event signal; or • passthrough steering module, which can control the remixing of media content and sound events to improve the playback experience. In particular, in some possible examples, the sound event detection and segmentation module may be used to detect the sound event type and its onset and offset. As will be understood and appreciated by the skilled person, this may be achieved by using any suitable means (e.g., signal processing and / or artificial intelligence (AI) -based techniques), depending on various implementations and / or circumstances. For instance, it may be suitably implemented, e.g., based on any one or more of the following techniques: • signal feature analysis and extraction cascading (e.g., in combination with) heuristic rules or decision tree; • signal feature analysis and extraction cascading machine learning, like Adaboost, XGboost, Gaussian mixture model (GMM), support vector machine (SVM), hidden Markov model (HMM), etc.; or • deep learning-based sound event detection, like convolutional neural network (CNN) model, etc. In some possible examples, the sound event detection and segmentation module may include any suitable mechanism to extract features in both frequency and time domain. For instance, for the spectrum features, features in a full band and different sub-bands may be calculated to capture pitch and harmonic characteristic differences. On the other hand, for the temporal features, short-term, medium-term and / or long-term features may be calculated to extract the time variation of the signal. Optionally, sound event enhancement module / algorithm may be used to suitably enhance the detected sound event signal. Depending on various implementations and / or circumstances, this may include digital signal processing (DSP) or AI -based background noise floor estimation and noise suppression, and the sound event detection result may be used to steer the background noise estimation. Further, the passthrough steering module / algorithm may be used to control the remixing of media content and sound events, thereby improving the playback experience. Depending on various implementations and / or circumstances, this may involve at least one of: • loudness analysis of sound event signal; • loudness analysis of playback media content; or • remix module / algorithm to blend sound event signal and playback media content. Reference is now to the figures, where the proposed techniques of the present disclosure are schematically shown. Therein, Fig.5 schematically illustrates an example block diagram 5000 of selective sound event passthrough according to some example embodiments of the present disclosure, where the input 5300 may be generally understood as an environment (or environmental) signal and the output 5800 may be generally understood as the blended media content 5600 and high quality sound events which are consistent with the real environment. As illustratively shown in Fig.6, the environmental signal 5300 may be captured by using any suitable means, such as a mono earbud outer microphone, a stereo earbud outer microphone, a headphone microphone, a smart glasses microphone, a mobile microphone, or the like. In particular, as shown in Fig.5, the proposed system architecture 5000 may comprise one or more of: a sound event type determination module 5200, a selective sound event detection and segmentation module 5400, an optional sound event enhancement module 5500, or a passthrough steering module 5700. In some possible examples, the specific sound event types may be decided by the user’s personalized selection (denoted as 5100 in Fig.5), for example, via a user interface (UI) as exemplarily shown in Fig.7. In some other possible example, the specific sound event types may be decided by the scene information, for example the target sound events may be car horn and siren on the street, or the like. Of course, as will be understood and appreciated by the skilled person, the sound event type may be determined or obtained by using any other suitable means, depending on various implementations and / or circumstances. For instance, the sound event types may be set by an application or by the mobile / wearable device (e.g., as a default setting, or the like). In some possible cases, for example when AI-based techniques are involved, the sound event types may be automatically determined (e.g., learned) during the course of classification or feature analysis. Incidentally, it may be worth mentioning that, even in the case of user selection for example via the UI, the list from which the sound event types are to be selected may not be always fixed, and may be extended on the way, for example by adding or recording (new / extra) sound event types that the user cares or is interested in now. Further, the sound event detection module / algorithm 5400 may be understood to aim to figure out what sound event is happening in the environment signal and when it is happening. The sound event type, and optionally, its onset and offset, may be detected. As briefly mentioned above, depending on various implementations and / or circumstances, the sound event detection can be realized (as illustratively shown in Fig.8) for example by utilizing at least one of: signal feature analysis and extraction cascading heuristic rules, signal feature analysis and extraction cascading machine learning, or deep learning-based sound event detection. Of course, as will be understood and appreciated by the skilled person, any other suitable means may be applied as well. In particular, in some possible examples, the feature extraction module (as illustratively shown in Fig.9) may be configured to extract the signal features, and thus determine whether there is a target sound event in the current frame. As will be understood and appreciated by the skilled person, for signal feature analysis and extraction, signal characteristics in frequency and / or time domain may be considered. For instance, for the spectrum features, features in a full band and / or different sub-bands may be calculated to capture pitch and harmonic characteristic differences. On the other hand, for the temporal features, short-term, medium-term, and / or long-term features may be investigated to extract the time variation of the signal. Some (non-limiting) examples of the possible discriminative features that are proposed to detect target sound events can be illustratively seen from Fig.9. In this example figure, it is assumed that the audio signal is first transformed into the frequency domain, and all of the features are calculated based on the frequency domain audio signal. In some possible examples, after feature extraction, a machine learning based model (for example, XG-Boost model, or the like) can be trained, as illustratively shown in Fig.10. Of course, as will be understood and appreciated by the skilled person, any other suitable machine learning based models may be adopted as well, depending on various implementations and / or circumstances. For instance, in some possible cases, firstly, the detector may realize polyphonic detection, which may be understood to be easy to add new sound event types, if necessary. In addition, the detector should output frame-based detection result. A corresponding block diagram for the inference is also schematically shown in Fig. 11, which will not be described in detail for the sake of conciseness. In some other possible examples, the feature extraction and sound event detection may be realized using deep learning based models (as illustratively shown in Figs.12 and 13), for example, a convolutional recurrent neural network (CRNN) model, or the like. Of course, as will be understood and appreciated by the skilled person, any other suitable machine learning based models may be adopted as well, depending on various implementations and / or circumstances. Similarly, the detector may realize polyphonic detection, and the detector should output frame-based detection result. In some possible cases, because some events may be from a noisy (e.g., outdoor) environment, the listening experience would be worse if the detected sound events are passed through directly. Therefore, in order to possibly guarantee the quality of the whole playback content, after the sound event detection module / algorithm 5400, an optional sound event enhancement module / algorithm 5500 could be added, in some possible implementations. Particularly, in some possible examples, the sound event enhancement module / algorithm 5500 may be used to remove the background noise overlapped with the sound event signal. As will be understood and appreciated by the skilled person, such enhancement may involve digital signal processing, DSP, and / or artificial intelligence, AI, -based techniques, or any other suitable techniques, depending on various implementations and / or circumstances. For instance, in some possible examples, the sound event enhancement module / algorithm 5500 may involve a DSP-based noise estimation cascading (e.g., in combination with) noise suppression, as illustratively shown in Fig.14. Particularly, in some possible examples, background (noise) estimation may be considered a power spectral estimation method based on optimal smoothing. In more detail, in some possible implementations, this may involve first calculating the smoothing parameter according to: '6K! = ^(LMN^:^O@^P^@P@QP^M^R@SNTP) (13)Then, the smoothed power spectral of the microphone signal may be calculated as: U= U ∗ '6K! + W^Q ∗ QM^X(W^Q) ∗ (1 − '6K!) (14)Thereafter, the process may involve searching for the minimum value U^>^in the time window; and calculating and updating the background power spectral accordingly: YZQ[\]MN^:GM^S@ = U^>^ (15)In some possible examples, the passthrough steering module / algorithm 5700 may be used to control the remixing of media content and sound events and improve the playback experience. As illustratively shown in Fig.15, depending on various implementations and / or circumstances, there may be, among others, at least one of the following possible factors to consider, namely: loudness analysis of sound event signal, loudness analysis of playback media content, or a remix module or mechanism to blend the sound event signal and the media content, in order to obtain the final output playback media content. Fig.16 is a flowchart of an example method 1600 for selective sound event passthrough for media content playback via a wearable device. The method 1600 may be performed by an electronic processor (for example, of the electronic apparatus 2800 of Fig.28), which may be configured to perform the method via execution of machine-executable instructions. The various process blocks illustrated in Fig.16 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. In particular, at step S1601, the method 1600 may detect a sound event in an external environment in reference to the wearable device. For example, the electronic processor 2801 may detect a sound event in an external environment in reference to the wearable device (e.g., wireless headphones). At step S1602, the method 1600 may determine sound event information corresponding to the detected sound event. For example, the electronic processor 2801 may determine sound event information (e.g., onset time, offset time, type of sound event, etc.) corresponding to the detected sound event. At step S1603, the method may enhance the detected sound event based on the determined sound event information. For example, the electronic processor 2801 may perform noise suppression on the detected sound event’s signal based on the determined sound event information (e.g., type). Finally, at step S1604, the method may remix the sound event and media content for playback via the wearable device. For example, the electronic processor 2801 may remix the sound event and media content for playback via the wearable device (e.g., wireless headphones). Fig.17 schematically illustrates another example of a method 1700 of selective sound event passthrough according to some example embodiments of the present disclosure. In particular, the method may be for selective sound event passthrough for a device (e.g., a wearable or mobile device) playing back media content (e.g., an audio, a video, or the like) in an (e.g., outdoor) environment. More particularly, the method 1700 may comprise, at step S1701, determining one or more target sound event types to be detected. The method 1700 may further comprise, at step S1702, detecting, in an environmental sound signal obtained from the environment, an environmental sound event matching any of the one or more target sound event types, and its respective onset and offset. Finally, the method 1700 may comprise, at step S1703, controlling remixing of the media content and the environmental sound event for playback via the device. To summarize the above, broadly speaking, at least some examples of the present disclosure may be seen to aim at providing an efficient, flexible, yet reliable mechanism for enabling / supporting selective sound event passthrough to pass the specific sound events (e.g., those considered of importance to the user, especially for safety issues, such as, a baby crying in living room, car-horn, siren alarms in the outdoor street, or the like) together with playback media content, and more particularly, to blend those real-world important sound events with consumer listening experience, thereby greatly improving the overall experience during use of such (wearable or mobile) devices. Accordingly, in some examples, a method and a system are presented for automatic detecting, segmentation, enhancement, and passthrough of selective sound event in real environment for portable electronic products to provide a tailored listening experience that contextually blends the real world with consumer media content. It may include any one or more of the following: • Sound event type determination module which can define the target sound event according to the user's personalized selection or adaptively. • Selective sound event detection and segmentation module which can figure out the sound event type and its onset and offset. • Sound event enhancement module which can remove background noise and get clean sound event signal. • Passthrough steering module which can control the remixing of media content and sound events to improve the playback experience. In some examples, the environmental signal can be captured by any one or more of: • MONO Earbud outer microphone, • Stereo Earbud outer microphone, • Headphone microphone, • Smart glasses microphone, or • Mobile microphone etc. In some examples, it is a device independent method; it can be deployed on mobile, ear buds, or any other device. In some examples, the specific sound event types can be decided by the user's personalized selection via UI. In some examples, the specific sound event types can be decided by the scene information analysis, for example the target sound events are car horn and siren on the street in adaptive mode. In some examples, the specific sound event types should be easily expanded. In some examples, the sound event detection and segmentation module will be used to detect the sound event type and its onset and offset. It can be: • Signal feature analysis and extraction cascading heuristic rules or decision tree. • Signal feature analysis and extraction cascading machine learning, like Adaboost, XGboost, GMM, SVM, HMM. • Deep learning-based sound event detection, like CNN model. In some examples, the sound event detection and segmentation module includes the method to extract features in both frequency and time domain. For the spectrum features, features in a full band and different sub-bands were calculated to capture pitch and harmonic characteristic differences. For the temporal features, medium-term and long-term features were calculated to extract the time variation of the signal. In some examples, the sound event enhancement algorithm will be used to enhance the detected sound event signal, which includes DSP-based background noise floor estimation and noise suppression. In some examples, the sound event detection result will steer the background noise estimation. In some examples, the passthrough steering algorithm will be used to control the remixing of media content and sound events and improve the playback experience, which includes: • Loudness analysis of sound event signal, • Loudness analysis of playback media content, • Remix module to blend sound event signal and playback media content. In some examples, the techniques described herein relate to a method for selective sound event passthrough during media content playback via a wearable device, the method including: detecting a sound event in an external environment in reference to the wearable device; determining sound event information corresponding to the detected sound event; enhancing the detected sound event based on the determined sound event information; and remixing the sound event and media content for playback via the wearable device. In some examples, the techniques described herein relate to a method, wherein the sound event information includes at least one of a sound event type, a sound event onset, and a sound event offset. In some examples, the techniques described herein relate to a method, wherein the wearable device includes one of headphones, earbuds, smart glasses, and / or head-mounted devices. In some examples, the techniques described herein relate to a method, further including defining a target sound event based on user selection or based on scene information analysis. In some examples, the techniques described herein relate to a method, wherein detecting a sound event in an external environment in reference to the wearable device includes capturing an external environment signal via a microphone. In some examples, the techniques described herein relate to a method, wherein the wearable device includes the microphone. In some examples, the techniques described herein relate to a method, wherein determining the sound event information corresponding to the detected sound event includes: extracting features of the sound event; and determining sound event information based on analysis of the extracted features. In some examples, the techniques described herein relate to a method, wherein the extracted features include features in the time domain and / or features in the frequency domain. In some examples, the techniques described herein relate to a method, wherein enhancing the detected sound event includes determining a background noise floor estimation of the detected sound event or performing noise suppression of the detected sound event. In some examples, the techniques described herein relate to a method, wherein remixing the sound event and media content for playback via the wearable device includes: determining loudness information of the enhanced sound event and loudness information of the media content; and based on the determined loudness information of the sound event and the media content, remixing the media content and the enhanced sound event. In some examples, the techniques described herein relate to a method, wherein the sound event includes a first audio signal and / or the media content includes a second audio signal. In some examples, the techniques described herein relate to a system for selective sound event passthrough during media content playback via a wearable device, the system including: a sound event detection module configured to: detect a sound event in an external environment in reference to the wearable device; determine sound event information corresponding to the detected sound event; a sound event enhancement module configured to: enhancing the detected sound event based on the determined sound event information; and a passthrough steering module configured to: remix the sound event and media content for playback via the wearable device. In some examples, the techniques described herein relate to a system, wherein the sound event detection module includes a machine learning model or a deep learning model. In some examples, the techniques described herein relate to an apparatus including a processor and a memory, configured to perform the method described above. In some examples, the techniques described herein relate to a computer program product including instructions which, when the program is executed by a computer, cause the computer to carry out the method described above. In some examples, the techniques described herein relate to a computer-readable storage medium storing the computer program product. Intelligent Conversation Awareness and Enhancement In many cases, people may like to use wearable audio devices like headsets and earbuds, to enjoy a personalized listening experience and attenuate the ambient sound. However, the user may not be able to hear the outside sound clearly to conduct a conversation unless a pass- through mode (of the wearable audio devices) is on (enabled), and the conversation is usually interfered by ambient noise, including, among others, interfering speech. In some possible examples, conversation awareness and enhancement may be used to automatically detect the user’s and other people’s speech in the conversation, so that the mobile audio can switch to pass-through mode intelligently, enhance the target speech during conversation, and return to playback mode after the conversation. Accordingly, in broad terms, the present disclosure also seeks to provide techniques for intelligent conversation awareness and enhancement. The goal of the proposed techniques may be understood as automatically detecting the time duration of conversation based on the signal captured by the microphones (and other suitable sensors, if available) on the wearable audio device, and enhancing the target speech and removing annoying intrusive noises during conversation mode. As discussed above, nowadays, people tend to use wearable audio devices, including earbuds and headphones, in a lot of ambiences and scenes, so that they can be isolated from the ambient sound. When the user wants to start a conversation with nearby people, they need to trigger pass-through mode and mute the playback media to hear the other people’s voices. Conversation awareness may be utilized to detect the time duration when the user has a conversation, so that the playback or noise cancellation can be adjusted to appropriate modes. There are typically microphones and sensors on the wearable audio devices, and the captured signal can be used to detect the speech in the conversation. Fig.18 shows an exemplary true wireless stereo (TWS) earbud microphone and speaker placement. In this example, there are two external microphones at the top and bottom of the stem to capture outside sound, an internal microphone inside the seal to capture the inside sound, and a micro speaker inside the seal for playback. Similarly, the headset and glasses usually also have microphones, loudspeakers, and other suitable sensors, so the proposed algorithm can also be utilized in those (wearable or mobile) devices. Still taking such a TWS earbud as an example, the ambient sound and the nearby people’s voice (sometimes also referred to as external speech) may be captured by the two external microphones, and the playback content may be played by the speaker and captured by the internal microphone. The user’s own speech (sometimes also referred to as self speech) may propagate in two ways. The first part may radiate from mouth to the air, then it can be captured by the external microphone. The second part may propagate for example by bone conduction and cause resonance in ear canal, then it is captured by the internal microphone. In the cases where a voice pickup unit (VPU) is available, the bone vibration of self speech may be captured by such VPU. Therefore, the self speech and external speech activity can be detected based on the microphone signal and the bone vibration signal. Furthermore, in some devices with a camera, such as smart glasses, the camera may also be configured to shoot video and images for external speech detection. Typically, the user (e.g., the wearer of the wearable device) is the one to start the conversation, and the conversation may be finished after a self speech or external speech segment. So, in some possible implementations, the conversation awareness may involve self speech detection, external speech detection, and turn-taking tracking. During the conversation, the external speech may usually be concatenated by the ambient sound, especially interference speech and intrusive noise, which may seriously affect the user’s conversation experience. Accordingly, in some possible implementations, conversation enhancement may be utilized to remove the annoying ambience noise and enhance the target speech during conversation. Based on these analyses, generally speaking, the present disclosure seeks to propose systems and methods to detect conversation and enhance conversation intelligently during mobile audio playback, more particularly, to detect and enhance the conversation between the wearable audio device user and the nearby people, thereby improving the user’s conversation experience. As will be discussed in more detail below, in broad terms, the proposed techniques can be implemented on devices with different configurations by fusing the short-term, mid-term, and / or long-term signal and information captured by the microphones, cameras, or other suitable sensors to estimate the conversation state, attenuate the interfering sound, and enhance target speech. For instance, in some possible examples, conversation awareness may be implemented as an all-in-one, parallel, series, or multi-stage system to keep low computational complexity and high reliability, and realized in signal processing, machine learning, or hybrid frame structures. As the proposed techniques mainly target mobile devices (such as wearable devices) and audio listening experience, some unique problems to solve concerning conversation awareness and enhancement may include: • detecting the conversation between the wearable device user and nearby people by: o detecting user’s self speech in various environments, o detecting external speech in the conversation, or o tracking conversation turn taking; and • optionally, enhancing the target speech in the conversation to improve the user’s listening and conversation experience. Reference is now made to Fig.19, which schematically illustrates an example block diagram of a conversation awareness system 1900 according to some example embodiments of the present disclosure. As shown in Fig.19, generally speaking, the proposed system (algorithm) architecture may be understood to comprise the following components, namely: a self speech detection module 1910, an external speech detection module 1920, a conversation turn-taking tracking module 1930, and, optionally in some possible examples, a conversation enhancement module 1940. As will be described in more detail below, in broad terms, the self speech detection module 1910 may be understood to detect the user’s self speech activity. The external speech detection module 1920 may be understood to detect other people’s voice(s) in the conversation. The conversation turn-taking tracking module 1930 may be understood to combine the information to track the conversation activity, and optionally, decide the onset, offset, and states of conversation. Finally, the optional conversation enhancement module 1940 may be understood to enhance the speech in the conversation and remove the annoying ambience noise, so that the listening experience can be improved. In some possible examples, the input of the self speech detection may include, among others, microphone signal, VPU signal, or the like. The input of external speech detection may include for example the self-speech detection result, the microphone signal, camera signal, and any other suitable sensor signal. Furthermore, the turn-taking tracking module may collect information from the other two modules to estimate the turn-taking of conversation, then decide if the conversation is dominated by self speech or external speech for example based on conversation patterns, and optionally, output the conversation state to inform the conversation enhancement (or, if applicable, any other suitable module / system). In some possible examples, the conversation enhancement module may attenuate the ambience sound and improve the conversation speech based on the signal from microphones, cameras, and any other suitable sensor, to reduce listening fatigue and improve the listening experience. In particular, as illustratively shown in Fig.20, the self speech detection module 1910 may be configured to detect the activity of self speech for example by analyzing the (internal and / or external) microphone signal, sensor signal, and the playback audio. It is to be noted that not all the above-illustrated input signals are necessary for self speech detection. On the other hand, as will also be understood and appreciated by the skilled person, any other suitable input may be fed to the self speech detection module 1910 as well, in order to facilitate the detection of the self speech activity. Particularly, based on various implementations of the hardware configuration and / or system framework, the self speech detection may have (but is certainly not limited to) the following various embodiments. Self speech detection with one internal and one external microphone internal and one external microphone, the self speech detection may be deployed as illustratively shown in Fig. 21. In detail, the internal microphone signal may be understood to contain both the self speech from the ear canal and the playback audio. An acoustic echo cancellation (AEC) module may be applied to the internal microphone signal, so that the playback audio component can be attenuated. As shown in Fig.21, the inner signal is dominated by self speech from the ear canal, but it may also be contaminated by the residual playback signal, because the AEC usually cannot remove the playback signal thoroughly. On the other hand, the outer signal from the external microphone may contain both the self speech and all the ambience sound, including other people’s speech, so its signal-to-noise ratio (SNR) is usually lower than the inner signal. The self-speech detection based on inner and outer signals may be understood to aim at detecting the speech segments by combining these two-channel (inner and outer) signals. As will be understood and appreciated by the skilled person, such self-speech detection can be implemented by involving any suitable machine learning (or more generally, AI-based), signal-processing, or hybrid algorithms / techniques. In some possible examples, the features for self speech detection may include time-domain, frequency domain, and / or time-frequency domain of the inner and outer signals. The comparison and combination of inner and outer signals may also be considered, because the self speech in inner and outer signals may be highly correlated. As for time and / or frequency domain features, the self speech may be different from the interference sound in signal-processing based features like amplitude, short-time time- frequency energy, envelope, zero-crossing rate, spectrum entropy, modulation spectrum, mel- frequency cepstrum, etc. Moreover, in some possible examples, self speech can also be differentiated from other sounds in auditory-based features like spectral patterns and gamma tone frequency cepstral coefficients, or the like. In some possible examples, the features can be applied to short, medium, or long duration to improve their reliability. In some possible examples, multi-resolution analysis based on multiple analysis window or tracking may also be considered useful to improve the reliability because the user naturally speak loud and long enough to make him or herself heard by nearby people, so that the features in different time-frequency window can be highly related with each other for the self speech, while the interference signal usually cannot meet such conditions. In some possible examples, speech-nonspeech differentiation may be applied to the inner signal, outer signal, or both. The self speech SNR in the inner signal is usually higher because it is less polluted by ambience sound, but its high-frequency components may usually be attenuated because of the bone conduction. On the other hand, the outer signal may have lower SNR, but the self speech is usually full-band. The combination of the results from two channels can be utilized to greatly improve the reliability of the self speech detection. In some possible cases, this may be considered especially important to deal with interference ambience speech, and AEC residual speech. Self speech detection with two or more external microphones on one side based on two external microphones on one TWS earbud or headset, the self speech detection may be deployed by detecting the speech based on the self speech components in the two microphone data. For example, if the earbud stem directs to the user’s mouth, as illustratively shown in Fig.22, the bottom external microphone is closer to the user’s mouth, so it has more self speech components than those in the top external microphone. Such configuration is may be considered to be relatively common for earbuds or headphones, so, in some possible examples, beamforming may also be applied to enhance self speech and attenuate the ambience sound. Then, in some possible examples, the features for speech detection can be further applied to detect the self speech onset and offset. Of course, as will also be understood and appreciated by the skilled person, any other suitable mechanisms / techniques that utilize the spatial selection ability brought out by such multi- microphone structure may be applied here as well. In some possible examples, based on the specific layout of the microphones, the beamforming may be integrated into the self speech detection or utilized explicitly before and / or after the speech detection. For instance, in some possible examples, if one external microphone is close enough to the mouth, then the detection can be based on one microphone, and the beamforming output can be used for confirmation, thereby improving reliability and reducing computation complexity. Self speech detection with external microphones on each side In some possible examples, if the self speech detection is deployed on a pair of microphones on the left and right sides of the user’s head, the binaural speech information may be applied to detect self speech. More particularly, because the two microphones both can receive speech signal from the user’s mouth, and the whole acoustic system is symmetric and stable, the binaural auditory features can be utilized to detect self speech. For example, interaural phase difference (IPD), interaural time difference (ITD), and interaural level difference (ILD) (or any other suitable metrics) of self speech may always be very close to zero, but the interaural correlation (IC) may always be high. Accordingly, the self speech components can be enhanced to detect the onset and offset. Self speech detection with only one microphone can only get signal from one external microphone, then techniques like voice identification / identifier (ID) or voice authentication may be exploited to detect the user’s self voice, so that the user’s speech can be differentiated from other sounds, especially other people’s speeches. In some possible examples, if the self speech detection can only get signal from one internal microphone, then techniques like exploiting an acoustic echo cancellation (AEC) module may be considered helpful to remove the playback components, and speech / nonspeech detection may be exploited to reliably detect the self speech. Self speech detection with VPU (e.g., bone conduction based) sensor (VPU) may be used to reliably detect the user’s self speech based on bone vibration. In some possible examples, if the wearable audio device is worn properly, it can be used to detect the voice segments of speech precisely. Depending on various implementations and / or circumstances, it may be utilized alone by itself or combined with other suitable acoustic signal from microphones to further improve the precision of voice segment(s) detection. Self speech detection on hardware of other configurations other suitable configuration of the microphones and / or sensors may be considered as the combination of the above illustrated examples. As will also be understood and appreciated by the skilled person, exploiting more sensors and / or microphones may very likely improve the reliability, for example by fusing the information in heuristic algorithms, machine learning models, or a hybrid of those. The self speech detection may be applied based on the configuration of hardware and fuse the raw signal and detection results in signal-processing, machine learning, or a hybrid thereof. Furthermore, potentially aiming at supporting the random conversation for users wearing the audio device, in some possible examples, the self speech detection may need to be kept always on. As a result, the global computational complexity may be considered important to keep practical battery efficiency. Accordingly, in some possible examples, the proposed techniques may be designed as a multi- stage detection framework, such that only the low complexity features and / or modules keep running. For example, in some possible implementations, only if the input signal has high energy and a nonstationary spectrum envelope, more complex features may then be calculated for further detection and confirmation. Next, techniques relating to external speech detection will be discussed. In broad terms, external speech generally means other people’s speech in the conversation, and the conversation typically alternates between self speech and external speech. In some possible cases, if only the conversation onset and offset are to be detected, external speech detection may be considered unnecessary when self speech is active. Generally speaking, external speech detection may be detected mainly based on the signals from microphones. Nevertheless, if there is a camera or any other suitable auxiliary information to exploit, the detection reliability may be improved. External speech detection with microphone(s) In some possible cases where external speech detection is performed with the help of microphones, the external speech may be detected based on its time-frequency and spatial features. As for time-frequency features, all suitable features to differentiate between speech and non-speech acoustic signals may be considered applicable. For instance, as for spatial features, the external speech is typically from the front side of the user, so that the user can have eye contact during the conversation. Naturally, the user would be unlikely to start a conversation with an unseen person to him or her. The user’s head and pinna amplify sound from the front, and they also have spatial filtering effect that helps to localize the sound sources from different directions. Accordingly, in some possible examples, based on the microphone configuration on the wearable hardware, the microphones may be utilized as a microphone array to enhance the signal from the target direction and suppress sound from the other side, then the external speech may be detected based on mono channel speech / non-speech detection. In some possible examples, if the microphones are on one side of the user, then beamforming may be considered a broadside or endfire microphone array. In some possible examples, if the synchronized microphones are on both the left and right sides of the user, then the binaural auditory based noise suppression may be utilized to enhance the sound from the target direction. To improve the computational efficiency of external speech detection, the modules may be combined into a multi-stage framework to detect speech from a target direction. For instance, in some possible examples as illustratively shown in Fig.23, the input microphone array signal may be processed by beamforming first, then speech / non-speech detection may be conducted on the enhanced signal. If the complexity of beamforming is low, such processing would be considered efficient. On the other hand, as illustratively shown in Fig.24, an initial speech detection may be conducted on a single-channel signal before beamforming, and the beamforming only starts after a possible speech signal is detected. Then, the final speech detection may be triggered to detect external speech. This processing chain may be understood to be able to save energy by reusing the self speech detection features on external microphone(s). As will be understood and appreciated by the skilled person, any other suitable multi-stage detection frame structure may be applied as well, for example by using low-complexity module(s) to trigger high complexity module(s), thereby keeping the global complexity low. Notably, in some possible examples, if the external speech could be limited to a group of known people, the speech may then be detected by the respective (e.g., preconfigured or predefined) voice ID of the people. Notably, such embodiment does not rely on the microphone number and configuration, but it can only work on known (e.g., registered) people. For instance, in some possible implementations, an online register may be applied if the person is not registered beforehand. External speech detection with camera(s) As mentioned above, in some possible examples, the external speech may be detected by camera signal, for example through detecting the talking person facing the user. Accordingly, in some possible examples, suitable techniques, such as face and / or lip movement detection, using one or more cameras, may be able to reliably detect the talking / conversation activity. Next, reference is made to Fig.25, which schematically illustrates an example of conversation turn-taking tracking according to some example embodiments of the present disclosure. In broad terms, conversation awareness may be generally understood to be able to detect at least one of the onset, offset, or activity of the conversation between the wearable audio device user and other people, so that the user can start a conversation by directly speaking some words, and the playback media and pass through setup can be changed automatically based on the self speech and external speech activity. Accordingly, conversation turn-taking tracking may be understood to be able to detect and track the conversation turn-taking status to inform the other related audio modules / systems on the device, as illustratively shown in Fig.25. Particularly, in a typical natural conversation, the dominant speaker may alternate between participants, and there may be states such as single talk (self speech, external speech), double talk, and mutual silence. After the self speech detection or external speech detection result changes, the state may jump to another state with certain probabilities, which may for example be preset based on statistical results or heuristic rules. The conversation turn-taking tracking can also be applied to other states to lower the computational complexity. For example, if the self speech detection is more reliable and always on, then external speech detection does not need to be kept working during self speech active state. If no self speech is detected, then the external speech can be triggered, so the system is in mutual silence mode, and both self speech and external speech detections are working to wait for the coming speech. In some possible examples, during this mutual silence state, the system may still be in conversation mode, but is counting a (e.g., preconfigured or predefined) waiting time (or hold on time). Depending on various implementations, the maximum mutual silence length (waiting time) may be heuristically set or adaptively updated, for example based on the history of speech burst lengths. For example, in some possible implementations, if the user keeps speaking long sentences, then the hold on time can be set a bit longer; whilst the short and quick conversation should have a relatively shorter hold on time. In some possible examples, during the hold on time, the conversation existence probability may be adjusted, for example according to signal(s) from microphones, cameras, sensors, or the like, so that the system can control the playback media playback and pass-through level to improve the user’s listening experience and shorten the waiting time. In some possible examples, conversation enhancement may be used to remove annoying ambience noise, and to enhance the conversation speech during the conversation. Since external speech is usually concatenated by the ambience sound, especially interference speech and intrusive noise, this may affect the conversation experience seriously. Aiming at attenuating the interference speech, which usually does not come from the target direction, speech enhancement based on for example any one or more of beamforming, face and lip movement, binaural auditory cues, or the like, may be considered possible approaches for speech enhancement and noise attenuation. Depending on various implementations and / or requirements, such techniques may be deployed as front-end or integrated with the conversation enhancement system. Moreover, as will be understood and appreciated by the skilled person, the system structure may be suitably based on signal-processing, machine learning (or more generally, AI based techniques), or a hybrid (combination) thereof. Furthermore, in some possible examples, the conversation enhancement may also be implemented for example by intelligently remixing the conversation with virtual media based on deep learning based speech isolation, or the like. Fig.26 is a flowchart of an example method 2600 for intelligent conversation awareness and enhancement for a wearable audio device. The method 2600 may be performed by an electronic processor 2801 (for example, of the electronic apparatus 2800 of Fig.28), which may be configured to perform the method via execution of machine-executable instructions. The various process blocks illustrated in Fig.26 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. In particular, at step S2601, the method 2600 may detect, via a self-speech detection module, speech of a user of the wearable audio device. For example, the electronic processor 2801 may detect, via the self-speech detection module, speech of the user of the wearable audio device (e.g., wireless headphones). At step S2602, the method 2600 may detect, via an external speech detection module, speech from an external source in reference to the user. For example, the electronic processor 2801 may detect, via the external speech detection module, speech from a person in conversation with the user. At step S2603, the method 2600 may track, via a conversation tracking module, a speech conversation between the user and the external source to determine conversation information. For example, the electronic processor 2801 may track, via the conversation tracking module, the speech conversation between the user and the person in conversation with the user to determine conversation information (e.g., conversation onset, conversation offset, who is speaking, the length of time each person speaks, conversation patterns, etc.). At step S2604, the method 2600 may determine, from the conversation information, one or more states of the conversation. For example, the electronic processor 2801 may determine, from the conversation information one or more states of the conversation (e.g., silent, active conversation, etc.). At step S2605, the method 2600 may, based on the one or more states, enhance the speech conversation in reference to content being played back on the wearable audio device. Fig.27 schematically illustrates another example of a method 2700 of conversation awareness according to some example embodiments of the present disclosure. In particular, the method 2700 may be for supporting conversation awareness during use of a device (e.g., a wearable or mobile device) of a user. More particularly, the method 2700 may comprise, at step S2701, detecting a self-speech activity of the user. The method 2700 may also comprise, at step S2702, detecting an external speech activity of one or more other participants of the conversation. Finally, the method 2700 may also comprise, at step S2703, tracking an activity of the conversation, based on the detection of the self-speech activity and / or the external speech activity. Configured as proposed, broadly speaking, the present disclosure generally provides an efficient, flexible, yet reliable mechanism for enabling / supporting intelligent conversation awareness, and potentially, also enhancement. For instance, as will be described in more detail below, this may involve, but certainly not limited to, being able to automatically detect the time duration of conversation based on the signal captured for example by microphone(s) and / or other suitable sensor(s) of the device, and optionally, to enhance the target speech and remove annoying intrusive noises during conversation mode. To summarize the above, broadly speaking, at least some examples of the present disclosure may be seen to aim at proposing an intelligent conversation awareness and enhancement system to detect and enhance the conversation between the wearable audio device user and the nearby people to improve the user’s conversation experience. It may be implemented on devices with different configurations by fusing the short-term, mid-term, and / or long-term signal and information captured by the microphones, cameras, and any other suitable sensors to estimate the conversation state, attenuate the interfering sound, and enhance target speech. Accordingly, in some examples, an intelligent conversation awareness and enhancement system and method is presented that includes: a. Conversation awareness with i. Self speech detection module, ii. External speech detection module, iii. Conversation turn-taking tracking module, b. Conversation enhancement module. In some examples, the conversation awareness algorithm is to automatically detect the user’s speech to find the conversation onset, then to track the user’s and other people’s speech in the conversation, and to detect the conversation offset. So that the mobile audio device can adjust the playback audio, pass through components, and enhance the target speech to support a better conversation and listening experience. In some examples, the self speech detection module is used to detect the user’s speech in cases when the user cannot hear the ambience sound. It may be applied to systems with various microphone and sensor configurations, including systems with a. Internal microphone and external microphone on one side, b. Two or more external microphones on one side, c. Synchronized external microphones on each side, d. Only one microphone, e. Voice pick up unit. In some examples, if the system has internal and external microphones, the playback audio in internal microphone signal needs to be attenuated to detect the self speech. The playback audio cancellation can be implemented as an acoustic echo cancellation module or integrated with the self speech detection. In some examples, if the system has internal and external microphones, then the self speech detection is based on the similarity of self speech components in internal and external microphone signals. The detection can be based on time domain, frequency domain, time- frequency domain features, and the comparison between the signals or features of the internal and external signal. In some examples, if the system has two or more external microphones on one side, beamforming can be utilized to enhance the user’s self speech and attenuate the ambience sound to improve self speech detection reliability. And it can be deployed before, after, or integrated to the self speech detection to improve the reliability. In some examples, if the system has synchronized external microphones on both sides, then binaural auditory features can be applied to enhance or detect self speech, and they can be integrated with other speech features in explicit or implicit ways. In some examples, if the system only has one microphone, then the user’s voiceID can be utilized to detect self speech. Because wearable audio devices are typically only used by constant users, voiceID is applicable for most devices. In some examples, if the system has voice pick up unit (VPU) to detect bone conducted voice or other sensors to detect user’s speaking action, self speech detection can be applied based on sensors. In some examples, self speech detection can use the signal from various microphones and sensors, and the information, signal, and features can be fused in heuristic, machine learning, or hybrid ways to improve reliability. The detection combines information from short, mid, and long time duration to improve reliability. In some examples, self speech detection can be deployed in various architectures to optimize the global computational complexity. The principle is to gate out most segments based on low complexity features or modules and confirm the detection with high complexity features or modules. In some examples, the external speech detection module is used to detect the person’s speech that the user is talking with. It may be applied to systems with various microphone and sensor configurations, including systems with a. Two or more microphones, b. Camera, c. Only one microphone. In some examples, the external speech detection utilizes speech-directivity features to detect the speech from target direction, in which the spatial filtering caused by pinna, head, and other physical structure is to be utilized. In some examples, if the system has microphones, external speech in the conversation can be detected based on voice activity detection for target directions. Beamforming with the microphones on one side or two sides can be integrated explicitly or implicitly to speech detection module to achieve directional selection for sound sources. In some examples, if the system has a camera, external speech detection can be based on the face or lip movement of the person in looking direction. In some examples, external speech detection can be deployed with voiceID to detect registered talker. In some examples, external speech detection can be deployed in various architectures to fuse the information from microphones, cameras, and sensors and optimize the global computational complexity. In some examples, the conversation turn-taking tracking module is to combine the self speech detection and external speech detection results to track the conversation activity, then decide the onset and offset or estimate the reliability of conversation. In some examples, the conversation turn-taking module estimates the conversation states based on duration and reliability of self speech detection and external speech detection, and the state transfer probability and waiting time length can be constant or adaptive based on heuristic or statistical rules. In some examples, the conversation enhancement module enhances the target speech and removes annoying intrusive noises during conversation mode. In some examples, the enhancement of target speech is by intelligently remixing the conversation with virtual media based on deep learning based speech isolation. In some examples, the removing of annoying intrusive noises can be based on signal processing based noise estimation and suppression or deep learning based noise suppression. In some examples, the techniques described herein relate to a method for intelligent conversation awareness and enhancement for a wearable audio device, the method including: detecting, via a self-speech detection module, speech of a user of the wearable audio device; detecting, via an external speech detection module, speech from an external source in reference to the user; tracking, via a conversation tracking module, a speech conversation between the user and the external source to determine conversation information; determining, from the conversation information, one or more states of the conversation; and based on the one or more states, enhancing the speech conversation to remove noise in content being played-back on the wearable audio device. In some examples, the techniques described herein relate to a method, wherein the detecting, via the self-speech detection module, speech of the user of the wearable audio device, further includes detecting speech of the user via one or more microphones. In some examples, the techniques described herein relate to a method, wherein the one or more microphones include at least one of an internal microphone of the wearable device, an external microphone of the wearable device, two or more external microphones of the wearable device, two synchronized external microphones of the wearable device, and / or a voice pick up unit of the wearable device. In some examples, the techniques described herein relate to a method, wherein the conversation information includes at least one of a conversation onset, a conversation offset, the speech of the user, and / or the speech from an external source. In some examples, the techniques described herein relate to a method, wherein detecting, via the self-speech detection module, speech of the user of the wearable audio device, includes: receiving an audio signal via the one or more microphones; and determining, based on a machine learning model or a deep learning model, that the audio signal includes the speech of the user. In some examples, the techniques described herein relate to an apparatus including a processor and a memory, configured to perform the method described above. In some examples, the techniques described herein relate to a computer program product including instructions which, when the program is executed by a computer, cause the computer to carry out the method described above. In some examples, the techniques described herein relate to a computer-readable storage medium storing the computer program product. Finally, Fig.28 is a schematic block diagram of an example electronic device or architecture 2800 suitable for implementing example embodiments of the present disclosure. Architecture 2800 may include, but is not limited to servers and client devices, systems, modules, and methods as described in reference to the preceding figures. As shown, the architecture 2800 includes central processing unit (CPU) 2801 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 2802 or a program loaded from, for example, storage unit 2808 to random access memory (RAM) 2803. The CPU 2801 may be, for example, an electronic processor 2801, which may include one or more processor cores, and in some examples, the processor 2801 may be multiple processors. In RAM 2803, the data used when CPU 2801 performs the various processes is also stored, as required. CPU 2801, ROM 2802, and RAM 2803 are connected to one another via bus 2804. Input / output (I / O) interface 2805 is also connected to bus 2804. The following components are connected to I / O interface 2805: input unit 2806, that may include a keyboard, a mouse, or the like; output unit 2807 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 2808 including a hard disk, or another suitable storage device; and communication unit 2809 which may include a network interface card such as a network card (e.g., wired or wireless). In some implementations, input unit 2806 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats). In some implementations, output unit 2809 includes systems with various numbers of speakers. Output unit 2807 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats). In some embodiments, communication unit 2809 is configured to communicate with other devices (e.g., via a network). Drive 2810 is also connected to I / O interface 2805, as required. Removable medium 2811, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 2810, so that a computer program read therefrom is installed into storage unit 2808, as required. A person skilled in the art would understand that although apparatus 2800 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alterations all fall within the scope of the present disclosure. In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 2809, and / or installed from the removable medium 2811, as shown in Fig.28. Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 2801 in combination with other components of Fig.28), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, a processor and / or other computing device(s), which may include control circuitry. While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. Interpretation device implementing the techniques described above can have the following example architecture. Other architectures are possible, including architectures with more or fewer components. In some implementations, the example architecture includes one or more processors (e.g., dual-core Intel® Xeon® Processors), one or more output devices (e.g., LCD), one or more network interfaces, one or more input devices (e.g., mouse, keyboard, touch-sensitive display) and one or more computer-readable mediums (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components can exchange communications and data over one or more communication channels (e.g., buses), which can utilize various hardware and software for facilitating the transfer of data and control signals between components. The term “computer-readable medium” refers to a medium that participates in providing instructions to processor for execution, including without limitation, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory) and transmission media. Transmission media includes, without limitation, coaxial cables, copper wire and fiber optics. Computer-readable medium can further include operating system (e.g., a Linux® operating system), network communication module, audio interface manager, audio processing manager and live content distributor. Operating system can be multi-user, multiprocessing, multitasking, multithreading, real time, etc. Operating system performs basic tasks, including but not limited to: recognizing input from and providing output to network interfaces and / or devices; keeping track and managing files and directories on computer-readable mediums (e.g., memory or a storage device); controlling peripheral devices; and managing traffic on the one or more communication channels. Network communications module includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols, such as TCP / IP, HTTP, etc.). Architecture can be implemented in a parallel processing or peer-to-peer infrastructure or on a single device with one or more processors. Software can include multiple software components or can be a single body of code. The described features can be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result. A computer program can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, a browser-based web application, or other unit suitable for use in a computing environment. Suitable processors for the execution of a program of instructions include, by way of example, both general and special purpose microprocessors, and the sole processor or one of multiple processors or cores, of any kind of computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer will also include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks; magneto- optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, ASICs (application-specific integrated circuits). To provide for interaction with a user, the features can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or a retina display device for displaying information to the user. The computer can have a touch surface input device (e.g., a touch screen) or a keyboard and a pointing device such as a mouse or a trackball by which the user can provide input to the computer. The computer can have a voice input device for receiving voice commands from the user. The features can be implemented in a computer system that includes a back-end component, such as a data server, or that includes a middleware component, such as an application server or an Internet server, or that includes a front-end component, such as a client computer having a graphical user interface or an Internet browser, or any combination of them. The components of the system can be connected by any form or medium of digital data communication such as a communication network. Examples of communication networks include, e.g., a LAN, a WAN, and the computers and networks forming the Internet. The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server. A system of one or more computers can be configured to perform particular actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions. While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination. Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the present disclosure discussions utilizing terms such as “processing”, “computing”, “calculating”, “determining”, “analyzing” or the like, refer to the action and / or processes of a computer or computing system, or similar electronic computing devices, that manipulate and / or transform data represented as physical, such as electronic, quantities into other data similarly represented as physical quantities. Reference throughout this disclosure to “one example embodiment”, “some example embodiments” or “an example embodiment” means that a particular feature, structure or characteristic described in connection with the example embodiment is included in at least one example embodiment of the present disclosure. Thus, appearances of the phrases “in one example embodiment”, “in some example embodiments” or “in an example embodiment” in various places throughout this disclosure are not necessarily all referring to the same example embodiment. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more example embodiments. As used herein, unless otherwise specified the use of the ordinal adjectives “first”, “second”, “third”, etc., to describe a common object, merely indicates that different instances of like objects are being referred to and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking, or in any other manner. Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted”, “connected”, “supported”, and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. In the claims below and the description herein, any one of the terms comprising, comprised of or which comprises is an open term that means including at least the elements / features that follow, but not excluding others. Thus, the term comprising, when used in the claims, should not be interpreted as being limitative to the means or elements or steps listed thereafter. For example, the scope of the expression a device comprising A and B should not be limited to devices consisting only of elements A and B. Any one of the terms including or which includes or that includes as used herein is also an open term that also means including at least the elements / features that follow the term, but not excluding others. Thus, including is synonymous with and means comprising. It should be appreciated that in the above description of example embodiments of the present disclosure, various features of the present disclosure are sometimes grouped together in a single example embodiment, Fig., or description thereof for the purpose of streamlining the present disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed example embodiment. Thus, the claims following the Description are hereby expressly incorporated into this Description, with each claim standing on its own as a separate example embodiment of this disclosure. Furthermore, while some example embodiments described herein include some but not other features included in other example embodiments, combinations of features of different example embodiments are meant to be within the scope of the present disclosure, and form different example embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed example embodiments can be used in any combination. In the description provided herein, numerous specific details are set forth. However, it is understood that example embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail in order not to obscure an understanding of this description. Thus, while there has been described what are believed to be the best modes of the present disclosure, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the present disclosure, and it is intended to claim all such changes and modifications as fall within the scope of the present disclosure. For example, any formulas given above are merely representative of procedures that may be used. Functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present disclosure. Enumerated Example Embodiments of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims. EEE 1. A context-aware noise compensation system used for headphones / earbuds, smart glasses, or head-mounted devices solves the playback media fluctuation problem caused by wearers’ vocal sounds, such as voice, coughing, sneezing, humming songs, etc., or rich transient sound events with large dynamic ranges in environmental sound, comprising: o Environmental sound perception via external microphones attached to headphones / earbuds, smart glasses, or head-mounted devices. o Context-aware information extraction contains wearer’s vocal sound detection, environmental sound event detection, or environmental sound source separation. o Context-aware noise compensation optimizes media in noisy environments to mask environmental sources to restore losses in loudness, naturalness, and timbre perception caused by environmental source masking. EEE 2. According to EEE 1, the external microphones capturing environment sound should be facing outward and closest to the ear canal to approximate human hearing. EEE 3. According to EEE 1, context-aware vocal sound detection can be implemented based on the Voice-Pick-Up (VPU) sensor data. For devices (some headphones, for example) without VPU attached, we can jointly leverage internal and external microphones to implement vocal sound detection. EEE 4. According to EEE 1, context-aware sound event detection can be implemented based on external microphone capture sound. Additionally, multiple external microphones can implement beamforming to enhance sounds from specific directions. EEE 5. According to EEE 1, context-aware environmental sound source separation can be implemented using a combination of VPUs and microphones. EEE 6. According to EEE 1, a context-aware noise compensation system is proposed to apply a time and frequency-dependent gain ^ to the playback media to mask the environmental source determined by the context-aware information and solve the playback media fluctuation problem. EEE 7. According to EEE 6, a time and frequency-dependent gain ^ is modeled to apply to the external microphone capture sound to represent the ratio of the environmental source to the whole capture sound. EEE 8. According to EEE 6, context-aware vocal sound detection ^^is binary value.If ^^ = 0, it corresponds to the whole external microphone capture sound at current frame asnoise to be masked. The excitation at this time frame will be updated based on ^ = 1. If ^^ =1, it corresponds to the external microphone capture sound at current frame is a mixture of vocal sound and environmental sound, then the excitation will keep the value of the pervious frame to avoid fluctuation caused by tracking the wearer’s vocal sound. EEE 9. According to EEE 6 and EEE 8, context-aware sound event detection ^7can indicate rich transient sound events with large dynamic ranges in the environment. Likewise, it can be used as ^^to mitigate playback media fluctuation caused by environmental sound events. EEE 10. According to EEE 6, context-aware environmental sound source separation can separate environmental sound sources. In this condition, ^ can be directly calculated by comparing the environmental sound source to the whole external microphone capture sound. An environmental sound source can be a noise floor, and then playback media can only mask the noise floor and ignore other sound events. Therefore, media fluctuation caused by rich transient sound events with large dynamic ranges in the environment can be mitigated. EEE 11. According to EEE 6, time and frequency-dependent optimized gain ^(^, ^^)can be derived resulting in G′!,^=G!H, e.g., the partial specific loudness of the media given the noisy environment being equal to the loudness of media in quiet. EEE 12. According to EEE 8, the vocal sound detection includes self-voice, coughing, sneezing, humming songs, etc. EEE 13. A context-aware noise compensation method for a wearable device, the method comprising: detecting one or more sounds using a microphone coupled to the wearable device; extracting contextual information corresponding to the one or more sounds; and based on the contextual information, performing noise compensation, via the wearable device. EEE 14. The method of EEE 13, wherein the one or more sounds comprise environmental sounds and / or vocal sounds. EEE 15. The method of EEE 14, wherein the wearable device comprises the microphone. EEE 16. The method of EEE 14 or EEE 15, wherein the detecting approximates human hearing. EEE 17. The method of EEE 16, wherein the microphone is configured to face outward and in proximity to an ear canal of a user of the wearable device, in reference to when the user wearing the wearable device. EEE 18. The method of any one of EEEs 13 to 17, wherein the wearable device comprises at least one of headphones, earbuds, smart glasses, and / or head-mounted devices. EEE 19. The method of any one of EEEs 13 to 18, wherein the contextual information comprises at least one of vocal sound detection information, environmental sound event detection information, and / or environmental source separation information. EEE 20. The method any one of EEEs 13 to 19, wherein the contextual information comprises vocal sound detection information, and wherein the vocal sound detection information comprises at least one of an indication of self-speech, coughing, sneezing, humming, and / or singing. EEE 21. The method of any one of EEEs 13 to 20, wherein the one or more sounds comprises a mixture of vocal sounds of a user of the wearable device and external environment sounds. EEE 22. The method of any one of EEEs 13 to 21, when the one or more sounds corresponds to a vocal sound, the detecting comprises detecting the one or more sounds based on Voice-Pick-Up sensor data or based on an internal microphone and an external microphone. EEE 23. The method of any one of EEEs 13 to 22, wherein the wearable device comprises a plurality of microphones, and wherein the detecting comprises beamforming the microphones to determine sounds from specific areas of the external environment. EEE 24. The method of any one of EEEs 13 to 23, wherein performing noise compensation, via the wearable device comprises: applying a gain to playback media of the wearable device to mask the at least a sound of the one or more sounds. EEE 25. The method of EEE 24, wherein the gain comprises a time and frequency- dependent gain. EEE 26. An apparatus comprising a processor and a memory, configured to perform the method according to any of EEEs 13 to 25. EEE 27. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method according to any of EEEs 13 to 25. EEE 28. A computer-readable storage medium storing the computer program product according to EEE 27. EEE 29. A method and system for automatic detecting, segmentation, enhancement and passthrough of selective sound event in real environment for portable electronic products to provide a tailored listening experience that contextually blends the real world with consumer media content. It is comprising of: o Sound event type determination module which can define the target sound event according to the user’s personalized selection or adaptively. o Selective sound event detection and segmentation module which can figure out the sound event type and its onset and offset. o Sound event enhancement module which can remove background noise and get clean sound event signal. o Passthrough steering module which can control the remixing of media content and sound events to improve the playback experience. EEE 30. The environmental signal can be captured by: o MONO Earbud outer microphone o Stereo Earbud outer microphone o Headphone microphone o Smart glasses microphone o Mobile microphone etc. It is a device independent method, it can be deployed on mobile, ear buds and other device. EEE 31. As in EEE 29, the specific sound event types can be decided by the user’s personalized selection via UI. EEE 32. As in EEE 29, the specific sound event types can be decided by the scene information analysis, for example the target sound events are car horn and siren on the street in adaptive mode. EEE 33. As in EEE 29, the specific sound event types should be easily expanded. EEE 34. As in EEE 29, the sound event detection and segmentation module will be used to detect the sound event type and its onset and offset, it can be: o Signal feature analysis and extraction cascading heuristic rules or decision tree. o Signal feature analysis and extraction cascading machine learning, like Adaboost, XGboost, GMM, SVM, HMM. o Deep learning-based sound event detection, like CNN model. EEE 35. As in EEE 34, the sound event detection and segmentation module includes the method to extract features in both frequency and time domain. For the spectrum features, features in a full band and different sub-bands were calculated to capture pitch and harmonic characteristic differences. For the temporal features, medium-term and long-term features were calculated to extract the time variation of the signal. EEE 36. As in EEE 29, the sound event enhancement algorithm will be used to enhance the detected sound event signal, it includes DSP-based background noise floor estimation and noise suppression. EEE 37. As in EEE 37, the sound event detection result will steer the background noise estimation. EEE 38. As in EEE 29, the passthrough steering algorithm will be used to control the remixing of media content and sound events and improve the playback experience, it includes: o Loudness analysis of sound event signal. o Loudness analysis of playback media content. o Remix module to blend sound event signal and playback media content. EEE 39. A method for selective sound event passthrough during media content playback via a wearable or mobile device, the method comprising: detecting a sound event in an external environment in reference to the wearable or mobile device; determining sound event information corresponding to the detected sound event; enhancing the detected sound event based on the determined sound event information; and remixing the sound event and media content for playback via the wearable or mobile device. EEE 40. The method of EEE 39, wherein the sound event information comprises at least one of a sound event type, a sound event onset, and a sound event offset. EEE 41. The method of EEE 39 or 40, wherein the wearable or mobile device comprises one of headphones, earbuds, smart glasses, mobile phone and / or head-mounted devices. EEE 42. The method of any one of EEEs 39 to 42, further comprising defining a target sound event based on user selection or based on scene information analysis. EEE 43. The method of any one of EEEs 39 to 42, wherein detecting a sound event in an external environment in reference to the wearable or mobile device comprises capturing an external environment signal via a microphone. EEE 44. The method of EEE 43, wherein the wearable device comprises the microphone. EEE 45. The method of any one of EEEs 39 to 44, wherein determining the sound event information corresponding to the detected sound event comprises: extracting features of the sound event; and determining sound event information based on analysis of the extracted features. EEE 46. The method of EEE 45, wherein the extracted features comprise features in the time domain and / or features in the frequency domain. EEE 47. The method of any one of EEEs 39 to 46, wherein enhancing the detected sound event comprises determining a background noise floor estimation of the detected sound event or performing noise suppression of the detected sound event. EEE 48. The method of any one of EEE 39 to 47, wherein remixing the sound event and media content for playback via the wearable device comprises: determining loudness information of the enhanced sound event and loudness information of the media content; and based on the determined loudness information of the sound event and the media content, remixing the media content and the enhanced sound event. EEE 49. The method of any one of EEEs 39 to 48, wherein the sound event comprises a first audio signal and / or the media content comprises a second audio signal. EEE 50. A system for selective sound event passthrough during media content playback via a wearable and / or mobile device, the system comprising: a sound event detection module configured to: detect a sound event in an external environment in reference to the wearable and / or mobile device; determine sound event information corresponding to the detected sound event; a sound event enhancement module configured to: enhancing the detected sound event based on the determined sound event information; and a passthrough steering module configured to: remix the sound event and media content for playback via the wearable device. EEE 51. The system of EEE 50, wherein the sound event detection module comprises a machine learning model or a deep learning model. EEE 52. An apparatus comprising a processor and a memory, configured to perform the method according to any of EEEs 39 to 49. EEE 53. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method according to any of EEEs 39 to 49. EEE 54. A computer-readable storage medium storing the computer program product according to EEE 53. EEE 55. An intelligent conversation awareness and enhancement algorithm comprising: a. Conversation awareness with i. Self speech detection module ii. External speech detection module iii. Conversation turn-taking tracking module b. Conversation enhancement module EEE 56. As in EEE 55, the conversation awareness algorithm is to automatically detect the user’s speech to find the conversation onset, then track the user’s and other people’s speech in the conversation and detect the conversation offset. So that the mobile audio device can adjust the playback audio, pass through components and enhance the target speech to support better conversation and listening experience. EEE 57. As in EEE 55, the self speech detection module is used to detect the user’s speech for the cases when the user cannot hear the ambience sound. It may be applied on system with various microphone and sensor configuration, including systems with a. Internal microphone and external microphone on one side b. Two or more external microphones on one side c. Synchronized external microphones in each side d. Only one microphone e. Voice pick up unit EEE 58. As in EEE 57, if the system has internal and external microphone, the playback audio in internal microphone signal needs to be attenuated to detect the self speech. The playback audio cancellation can be implemented as an acoustic echo cancellation module or be integrated with the self speech detection. EEE 59. As in EEE 58, if the system has internal and external microphone, then the self speech detection is based on the similarity of self speech components in internal and external microphone signal. The detection can be based on time domain, frequency domain, time-frequency domain features and the comparison between the signals or features of the internal and external signal. EEE 60. As in EEE 57, if the system has two or more external microphones on one side, beamforming can be utilized to enhance the user’s self speech and attenuate the ambience sound to improve self speech detection reliability. And it can be deployed before, after or integrated to the self speech detection to improve the reliability. EEE 61. As in EEE 57, if the system has synchronized external microphones on both sides, then binaural auditory features can be applied to enhance or detect self speech, and they can be integrated with other speech features in explicit or implicit ways. EEE 62. As in EEE 57, if the system only has one microphone, then the user’s voice ID can be utilized to detect self speech. Because wearable audio devices are typically only used by constant users, voice ID is applicable for most devices. EEE 63. As in EEE 57, if the system has voice pick up unit (VPU) to detect bone conducted voice or other sensors to detect user’s speaking action, self speech detection can be applied based on sensors. EEE 64. As in EEE 57, self speech detection can use the signal from various microphones and sensors, and the information, signal and features can be fused in heuristic, machine learning or hybrid ways to improve reliability. The detection combines information from short, mid and long time duration to improve reliability. EEE 65. As in EEE 57, self speech detection can be deployed in various architectures to optimize the global computational complexity. The principle is to gate out most segments based on low complexity features or modules and confirm the detection with high complexity features or modules. EEE 66. As in EEE 55, the external speech detection module is used to detect the person’s speech that the user is talking with. It may be applied on system with various microphone and sensor configuration, including systems with a. Two or more microphones b. Camera c. Only one microphone EEE 67. As in EEE 56, the external speech detection utilizes speech-directivity features to detect the speech from target direction, in which the spatial filtering caused by pinna, head and other physical structure are to be utilized. EEE 68. As in EEE 66, if the system has microphones, external speech in the conversation can be detected based on voice activity detection for target directions. Beamforming with the microphones on one side or two sides can be integrated explicitly or implicitly to speech detection module to achieve directional selection for sound sources. EEE 69. As in EEE 66, if the system has camera, external speech detection can be based on the face or lip movement of the person in looking direction. EEE 70. As in EEE 66, external speech detection can be deployed with voice ID to detect registered talker. EEE 71. As in EEE 66, external speech detection can be deployed in various architectures to fuse the information from microphones, cameras and sensors and optimize the global computational complexity. EEE 72. As in EEE 55, the conversation turn-taking tracking module is to combine the self speech detection and external speech detection results to track the conversation activity, then decide the onset and offset or estimate the reliability of conversation. EEE 73. As is in EEE 72, the conversation turn-taking module estimates the conversation states based on duration and reliability of self speech detection and external speech detection, and the state transfer probability and waiting time length can be constant or adaptive based on the heuristic or statistical rules. EEE 74. As in EEE 55, the conversation enhancement module enhances the target speech and remove annoying intrusive noises during conversation mode. EEE 75. As in EEE 74, the enhancement of target speech is by intelligent remixing the conversation with virtual media based on deep learning based speech isolation. EEE 76. As in EEE 74, the removing of annoying intrusive noises can be based on signal processing based noise estimation and suppression or deep learning based noise suppression. EEE 77. A method for intelligent conversation awareness and enhancement for a wearable audio device, the method comprising: detecting, via a self-speech detection module, speech of a user of the wearable audio device; detecting, via an external speech detection module, speech from an external source in reference to the user; tracking, via a conversation tracking module, a speech conversation between the user and the external source to determine conversation information; determining, from the conversation information, one or more states of the conversation; and based on the one or more states, enhancing the speech conversation in reference to content being played-back on the wearable audio device. EEE 78. The method of EEE 77, wherein the detecting, via the self-speech detection module, speech of the user of the wearable audio device, further comprises detecting speech of the user via one or more microphones. EEE 79. The method of EEE 78, wherein the one or more microphones comprise at least one of an internal microphone of the wearable device, an external microphone of the wearable device, two or more external microphones of the wearable device, two synchronized external microphones of the wearable device, and / or a voice pick up unit of the wearable device. EEE 80. The method of any one of EEEs 77 to 79, wherein the conversation information comprises at least one of a conversation onset, a conversation offset, the speech of the user, and / or the speech from an external source. EEE 81. The method of any one of EEEs 77 to 80, wherein the one or more states of conversation comprise at least one of mutual silence, self speech, external speech, and / or double speech. EEE 82. The method of any one of EEEs 78 to 81 when dependent on EEE 78, wherein detecting, via the self-speech detection module, speech of the user of the wearable audio device, comprises: receiving an audio signal via the one or more microphones; and determining, based on a machine learning model or a deep learning model, that the audio signal comprises the speech of the user. EEE 83. The method of any one of EEE 77 to 82, wherein enhancing the speech conversation comprises removing ambient noise and / or enhancing the conversation speech during a conversation state of the one or more states of the conversation. EEE 84. The method of EEE 83, wherein enhancing the conversation speech comprises performing low complexity noise estimation and noise suppression. EEE 85. The method of EEE 83, wherein enhancing the conversation speech comprises performing deep learning based speech separation or isolation and intelligent remixing conversation, ambience and virtual media. EEE 86. The method of EEE 83, wherein enhancing the conversation speech comprises beamforming via multiple microphones or binaural auditory cures. EEE 87. The method of EEE 83, wherein enhancing the conversation speech comprises performing audio / video beamforming based on multi-modal information. EEE 88. The method of EEE 87, wherein the multi-modal information comprises face and lip information. EEE 89. An apparatus comprising a processor and a memory, configured to perform the method according to any of EEEs 77 to 88. EEE 90. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method according to any of EEEs 77 to 88. EEE 91. A computer-readable storage medium storing the computer program product according to EEE 90.

Claims

CLAIMS 1. A context-aware noise compensation system for a wearable device playing back media content to a user in a noisy environment, wherein the wearable device comprises one or more sensors including: acoustic sensors such as microphones, or non-acoustic sensors such as accelerometers, and wherein the system is configured to: obtain a sound signal from the noisy environment; determine context information associated with the sound signal, wherein the context information comprises at least one of: vocal sound detection information indicative of presence of one or more vocal sounds of the user in the sound signal, or sound event detection information indicative of presence of one or more environmental sound events in the sound signal; and optimize, based on the context information, the media content for playback by the wearable device.

2. The system according to claim 1, wherein the one or more vocal sounds of the user comprise self-voice, such as speech, coughing, sneezing, or humming, of the user.

3. The system according to claim 1 or 2, wherein the one or more environmental sound events comprise at least one transient sound event with a large dynamic range in the environment.

4. The system according to any one of the preceding claims, wherein the wearable device includes: a headphone, an earbud, a smart glass, or a head-mounted device, HMD.

5. The system according to any one of the preceding claims, wherein the sound signal is obtained by using one or more external microphones of the wearable device.

6. The system according to any one of the preceding claims, wherein the vocal sound detection information is determined based on internal and external microphones of the wearable device jointly.

7. The system according to any one of the preceding claims, wherein the one or more sensors comprise a voice-pick-up, VPU, sensor; andwherein the vocal sound detection information is determined based on the VPU sensor.

8. The system according to claim 7, wherein the context information further comprises information indicative of one or more sound sources separated by using the VPU sensor and the microphones.

9. The system according to any one of the preceding claims, wherein the sound event detection information is determined based on one or more external microphones of the wearable device.

10. The system according to any one of the preceding claims, wherein the optimization of the media content involves: if one or more vocal sounds are detected or the detected one or more vocal sounds exceed a predetermined threshold level, estimating a current level of the environment based on microphone data from the past; otherwise, determining the current level of the environment based on current microphone data.

11. The system according to any one of the preceding claims, wherein the optimization of the media content involves: determining, based on the vocal sound detection information, a ratio of environmental sound sources to the sound signal.

12. The system according to claim 11, wherein if the vocal sound detection information indicates absence of vocal sounds at a current frame, the ratio is set to 1 for updating an excitation level of the environmental sound sources at the current frame; and if the vocal sound detection information indicates presence of one or more vocal sounds at the current frame, the ratio is determined so as to keep the excitation level of the previous frame.

13. The system according to claim 11 or 12, wherein the ratio ^ is determined according to: 1,^^ ^ = 0^^ ^1,where ^^ represents the vocal sound detection information, and ^^(^, ^^) denotes anexcitation level of the environmental sound sources at an ^-th frame at a center frequency ^^.

14. The system according to claim 11 when depending on claim 8, wherein the ratio is determined so as to mask a specific environmental sound source that has been separated.

15. The system according to claim 14, wherein the ratio ^ is determined according to: ^(^, ^ ^ (^,^^) = ^ ^ ^)^^(^,^^),where ^^(^, ^^) denotes an sound sources at an^-th frame at a center frequency ^ ,^the separated environmentalsound source.

16. The system according to any one of the preceding claims, wherein the optimization of the media content involves applying a time frequency dependent gain to the media content.

17. The system according to claim 16, wherein the gain is determined such that, when being applied, a partial specific loudness of the media content in the noisy environment is equal to a specific loudness for normal hearing and in quiet.

18. The system according to claim 16 or 17, wherein the gain ^ is determined such that the following equation is fulfilled: ^^(^^ ∙ ^! + # + ^^ ∙ ^^)$ − (# + ^^ ∙ ^^)$& = ^^(^! + #)$ − #$&,where ^, #, and ' are constants, ^!denotes a running short-term estimate of excitation of the media content in quiet, ^ denotes a ratio of environmental sound sources to the sound signal, and ^^denotes an excitation level of the environmental sound sources.

19. The system according to claim 18, wherein the gain ^ is determined according to: 0 / ,^^ = ((^)*+),^+,*(+*-.∙^^), / ^+*-.∙^^ .

20. The system according to any one of the preceding claims, wherein the determination of the context information involves signal processing and / or artificial intelligence, AI, based techniques.

21. A method of context-aware noise compensation for a wearable device playing back media content to a user in a noisy environment, the method comprising: obtaining a sound signal from the noisy environment; determining context information associated with the sound signal, wherein the context information comprises at least one of: vocal sound detection information indicative of presence of one or more vocal sounds of the user in the sound signal, or sound event detection information indicative of presence of one or more environmental sound events in the sound signal; and optimizing, based on the context information, the media content for playback by the wearable device.

22. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the apparatus to carry out the method according to claim 21.

23. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to claim 21.

24. A computer-readable storage medium storing the program according to claim 23.

25. A system for selective sound event passthrough for a device playing back media content in an environment, the system comprising: a sound event type determination module configured for determining one or more target sound event types to be detected; a sound event detection and segmentation module configured for detecting, in an environmental sound signal obtained from the environment, an environmental sound event matching any of the one or more target sound event types, and its respective onset and offset; and a passthrough steering module configured for controlling remixing of the media content and the environmental sound event for playback via the device.

26. The system according to claim 25, wherein the sound event type determination module is configured for determining the one or more target sound event types to be detected based on user selection, pre- or default configuration, or automatically.

27. The system according to claim 25 or 26, wherein the environmental sound signal is obtained by using one or more microphones of the device.

28. The system according to any one of claims 25 to 27, wherein the detection of the environmental sound event and the respective onset and offset thereof involves rule and / or artificial intelligence, AI, -based techniques.

29. The system according to any one of claims 25 to 28, wherein the detection of the environmental sound event and the respective onset and offset thereof involves feature extraction of the environmental sound signal.

30. The system according to any one of claims 25 to 29, wherein the feature extraction involves: extracting full and / or subband features in a frequency domain, thereby capturing pitch and harmonic characteristic differences of the environmental sound signal; and / or extracting short and / or long -term temporal features in a time domain, thereby extracting time variation of the environmental sound signal.

31. The system according to any one of claims 25 to 30, wherein the passthrough steering module is specifically configured to perform, for the controlling of the remixing of the media content and the environmental sound event: loudness analysis of the environmental sound event; and loudness analysis of the media content.

32. The system according to any one of claims 25 to 31, wherein the system further comprises a sound event enhancement module configured for enhancing the environmental sound event, for the remixing.

33. The system according to claim 32, wherein the enhancement of the environmental sound event involves digital signal processing, DSP, and / or artificial intelligence, AI, -based techniques.

34. The system according to claim 32 or 33, wherein the enhancement of the environmental sound event involves noise estimation and noise suppression.

35. The system according to claim 34, wherein the noise estimation involves: calculating a smoothing parameter based on the environmental sound event; calculating a smoothed power spectral of the environmental sound signal; searching for a minimum value of the smoothed power spectral in a time window; and calculating and updating the noise estimation based on the minimum value.

36. The system according to any one of claims 25 to 35, wherein the device is a mobile device, or a wearable device such as: a headphone, an earbud, a smart glass, or a head- mounted device, HMD.

37. The system according to any one of claims 25 to 36, wherein the one or more target sound event types to be detected comprises at least one of: siren alarm, baby crying, car horn, door knocking, or house alarm.

38. A method of selective sound event passthrough for a device playing back media content in an environment, the method comprising: determining one or more target sound event types to be detected; detecting, in an environmental sound signal obtained from the environment, an environmental sound event matching any of the one or more target sound event types, and its respective onset and offset; and controlling remixing of the media content and the environmental sound event for playback via the device.

39. The method according to claim 38, wherein the method further comprises: obtaining the environmental sound signal by using one or more microphones of the device.

40. The method according to claim 38 or 39, wherein the method further comprises, before the remixing: enhancing the environmental sound event.

41. The method according to any one of claims 38 to 40, wherein the method involves one or more artificial intelligence, AI, based techniques.

42. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the apparatus to carry out the method according to any one of claims 38 to 41.

43. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 38 to 41.

44. A computer-readable storage medium storing the program according to claim 43.

45. A system for supporting conversation awareness during use of a device of a user, the system comprising: a self-speech detection module configured for detecting a self-speech activity of the user; an external speech detection module configured for detecting an external speech activity of one or more other participants of the conversation; and a conversation turn-taking tracking module configured for, based on the detection of the self-speech activity and / or the external speech activity, tracking an activity of the conversation.

46. The system according to claim 45, wherein the self-speech activity of the user is detected based on at least one of: a microphone signal of the device, a sensor signal of the device, or a media signal played back by the device.

47. The system according to claim 45 or 46, wherein the device comprises: an internal microphone capable of capturing self-speech of the user, thereby obtaining an internal microphone signal; and an external microphone capable of capturing self-speech of the user, thereby obtaining an external microphone signal that further comprises ambient sound; and wherein the self-speech activity of the user is detected based on the internal microphone signal and the external microphone signal.

48. The system according to claim 47, wherein the internal microphone signal further comprises a media signal played back by the device; and wherein acoustic echo cancellation, AEC, is applied to the internal microphone signal for suppressing or attenuating the media signal, before the internal microphone signal is used for the detection of the self-speech activity of the user.

49. The system according to claim 47 or 48, wherein the detection of the self-speech activity of the user involves digital signal processing, DSP, and / or artificial intelligence, AI, - based techniques.

50. The system according to claim 45 or 46, wherein the device comprises: two or more external microphones on one side of the user, each configured for obtaining a respective external microphone signal that comprises self-speech of the user and ambient sound; and wherein the detection of the self-speech activity of the user involves one or more spatial selection-based techniques, such as beamforming.

51. The system according to claim 45 or 46, wherein the device comprises: a respective external microphone on each side relative to the user, each configured for obtaining a respective external microphone signal that comprises self-speech of the user and ambient sound; and wherein the self-speech activity of the user is detected based on binaural speech information derived from the external microphone signals.

52. The system according to claim 51, wherein the binaural speech information comprises: interaural phase difference, IPD, interaural time difference, ITD, and / or interaural level difference, ILD.

53. The system according to claim 45 or 46, wherein the device comprises one microphone; wherein if the microphone is an external microphone configured for obtaining an external microphone signal that comprises self-speech of the user and ambient sound, the detection of the self-speech activity of the user involves voice identification and / or voice authentication; andwherein if the microphone is an internal microphone configured for obtaining an internal microphone signal that comprises self-speech of the user and a media signal played back by the device, the detection of the self-speech activity of the user involves acoustic echo cancellation, AEC, for suppressing or attenuating the media signal.

54. The system according to any one of claims 45 to 53, wherein the device comprises a voice pick-up, VPU, sensor; and wherein the self-speech activity of the user is detected based on data obtained by the VPU sensor.

55. The system according to any one of claims 45 to 54, wherein the external speech activity is detected based on signals obtained from one or more microphones of the device.

56. The system according to claim 55, wherein the detection of the external speech activity involves extracting at least one of: time, frequency, or spatial, features from the signals.

57. The system according to claim 55 or 56, wherein the device comprises a microphone array; and wherein the external speech activity is detected based on a configuration or layout of the microphone array.

58. The system according to claim 57, wherein if the microphone array comprises two or more microphones on one side of the user, the detection of the external speech activity involves beamforming; and if the microphone array comprises a respective microphone on each side relative to the user, the detection of the external speech activity involves binaural auditory based noise suppression for enhancing sound from a target direction but suppressing sounds from other directions.

59. The system according to claim 58, wherein a speech detection is performed on a single channel signal obtained by one microphone of the microphone array, before triggering the beamforming.

60. The system according to any one of claims 55 to 59, wherein the device comprises a camera; and wherein the external speech activity is detected further based on face and / or lip movement detection by using the camera.

61. The system according to any one of claims 45 to 60, wherein the tracking of the activity of the conversation involves determining at least one of: onset, offset, or states of the conversation.

62. The system according to claim 61, wherein the states of the conversation comprise: a self-speech state, an external speech state, a double talk state, and a mutual silence state.

63. The system according to claim 61 or 62, wherein transition between the states of the conversation is based on the detection of the self-speech activity and / or the external speech activity.

64. The system according to claim 63, wherein the transition is further based on a statistic or heuristic rule -based probability.

65. The system according to any one of claims 62 to 64, wherein the mutual silence state is configured with a hold on timer that is heuristically or adaptively configurable.

66. The system according to any one of claims 45 to 65, wherein the system further comprises a conversation enhancement module configured for enhancing the conversation.

67. The system according to claim 66, wherein the enhancement of the conversation involves at least one of: enhancing speech during the conversation, suppressing ambience or background noise, attenuating interference speech, or stop / replay of media content played back by the device.

68. The system according to claim 66 or 67, wherein the enhancement of the conversation involves remixing the conversation with virtual media based on deep learning based speech isolation.

69. The system according to any one of claims 66 to 68, wherein the enhancement of the conversation involves digital signal processing, DSP, and / or artificial intelligence, AI, - based techniques.

70. The system according to any one of claims 45 to 69, wherein the device is a mobile device, or a wearable device such as: a headphone, an earbud, a smart glass, or a head- mounted device, HMD.

71. A method of supporting conversation awareness during use of a device of a user, the method comprising: detecting a self-speech activity of the user; detecting an external speech activity of one or more other participants of the conversation; and tracking an activity of the conversation, based on the detection of the self-speech activity and / or the external speech activity.

72. The method according to claim 71, wherein the method further comprises enhancing the conversation.

73. The method according to claim 71 or 72, wherein the method involves one or more artificial intelligence, AI, -based techniques.

74. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the apparatus to carry out the method according to any one of claims 71 to 73.

75. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 71 to 73.

76. A computer-readable storage medium storing the program according to claim 75.

Citation Information

Patent Citations

  • Ambient sound enhancement and acoustic noise cancellation based on context

    EP3745736A1

  • A hearing aid comprising an own voice conversation tracker

    EP3930346A1

  • Voice-Enhanced Awareness Mode

    US20170194020A1

  • Apparatus and method for operating wearable device

    US20210407532A1

  • Method for controlling ambient sound and electronic device therefor

    US20220189477A1