Controlling output of audio data

Through a machine learning model based on the user's neural activity, it identifies the user's auditory attention and automatically adjusts the audio output in the headphone device, solving the problem of users switching between different audio sources and realizing dynamic audio control and intelligent audio management of augmented reality headphones.

CN120708651APending Publication Date: 2025-09-26NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317808.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

It is difficult to automatically manage users' switching needs for different audio sources at different times, and existing technologies find it difficult to effectively identify users' auditory attention and adjust audio output accordingly.

Method used

By using machine learning models to identify the user's auditory attention based on measurements of the user's neural activity, and by controlling the combination of speakers and microphones, the output of the audio dataset, including audio sources for audio tracks and real-world audio scenes, is automatically adjusted to achieve dynamic switching and audio control in augmented reality headset devices.

Benefits of technology

It achieves automatic audio output adjustment based on user attention, improves user experience, saves energy and does not require direct user interaction, can switch between ANC and transparency modes, and enhances the adaptability of audio datasets and users' perception of real-world audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708651A_ABST
    Figure CN120708651A_ABST
Patent Text Reader

Abstract

Example embodiments relate to an apparatus, method, and computer program product relating to controlling output of audio data. An example method is disclosed that includes outputting a first set of audio data via one or more speakers; capturing the real-world audio scene via the one or more microphones to provide a second audio data set for output via the one or more speakers; and identifying which of the first audio data set and at least a portion of the real-world audio scene has auditory attention of the user based on the measured neural activity of the user. The method may also include controlling output of at least some of the first and / or second audio data sets via the one or more speakers based on the identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various example embodiments are directed to controlling the output of audio data, for example, based on measured neural activity of a user. Background Art

[0002] A user may be interested in different audio sources at different times. For example, a user may be listening to an audio track via the speakers of a headphone device, but at one or more times, may also realize that they may want to hear the real-world audio sounds around them, at least temporarily, rather than the audio track. Knowing which audio has the user's auditory attention at a given time may be useful. Summary of the Invention

[0003] The scope of protection sought by various embodiments of the present invention is defined by the independent claims. Embodiments and features described in this specification that do not fall within the scope of the independent claims (if any) are to be construed as examples useful for understanding various embodiments of the present invention.

[0004] According to a first aspect, an apparatus is described comprising: means for outputting a first audio dataset via one or more speakers of the apparatus; means for capturing a real-world audio scene external to the apparatus via one or more microphones of the apparatus to provide a second audio dataset for output via the one or more speakers; means for identifying which of the first audio dataset and at least a portion of the real-world audio scene has the user's auditory attention based on measured neural activity of the user; and means for controlling output of at least some of the first and / or second audio datasets via the one or more speakers based on the identification.

[0005] In some example embodiments, the first audio data set may represent an audio track or a communication session received from a user device associated with the apparatus.

[0006] In some example embodiments, in response to identifying that the first audio data set has the user's auditory attention, the means for controlling the output may be configured to disable or attenuate output of the second audio data set via the one or more speakers.

[0007] In some example embodiments, the apparatus may further include: means for generating a noise cancellation signal based on the captured real-world audio scene, wherein the means for controlling the output is configured to: further in response to identifying that the first audio data set has the user's auditory attention, enable or increase a gain associated with the noise cancellation signal for output to the one or more speakers.

[0008] In some example embodiments, the first audio data set may represent multiple audio sources, wherein the means for identifying may be configured to: identify that a first audio source among the multiple audio sources has the user's auditory attention; and the means for controlling the output may be configured to: amplify the first audio source relative to the other(s) audio sources.

[0009] In some example embodiments, in response to identifying that at least a portion of the real-world audio scene has the user's auditory attention, the means for controlling the output may be configured to disable or reduce a gain associated with the first audio data set.

[0010] In some example embodiments, in response to identifying that at least a portion of the real-world scene has the user's auditory attention, the component for controlling the output may be configured to: enable or increase the gain associated with at least some of the second audio data sets for output to the one or more speakers.

[0011] In some example embodiments, the means for controlling the output may be configured to: further in response to identifying that at least a portion of the real-world audio scene has the user's auditory attention, disable or reduce a gain associated with the noise cancellation signal for output to the one or more speakers.

[0012] In some example embodiments, a real-world audio scene may include multiple real-world audio sources, the component for identifying may be configured to: identify that a first real-world audio source among the multiple real-world audio sources has the user's auditory attention, and the device may further include: a component for steering a sound capture beam of one or more microphones in the direction of the first real-world audio source so that the audio signal of the first real-world audio source is output to the one or more speakers with a higher gain than other (one or more) real-world audio sources.

[0013] In some example embodiments, the apparatus may further include: means for dividing the second audio data set into a plurality of frequency sub-bands, wherein the means for identifying may be configured to: identify that a first frequency sub-band among the plurality of frequency sub-bands has the user's auditory attention, and the means for controlling the output may be configured to: amplify the output of the first frequency sub-band with a higher gain than other frequency sub-bands.

[0014] In some example embodiments, the means for controlling the output may be configured to, further in response to identifying that the real-world audio scene has the user's auditory attention, disable or reduce a gain associated with a noise cancellation signal corresponding to a first frequency sub-band. In some example embodiments, the first frequency sub-band may correspond to speech audio.

[0015] In some example embodiments, the apparatus may further comprise: a component for measuring a user's neural activity. In some example embodiments, the apparatus may be comprised by a headphone device.

[0016] In some example embodiments, the apparatus may include active noise cancellation functionality operable in a transparent mode for outputting at least some of the second audio data set via one or more speakers, and the components for controlling the output of at least some of the second audio data set may be configured to at least enable or control a gain at least associated with the transparent mode.

[0017] According to a second aspect, a method is described, comprising: outputting a first audio dataset via one or more speakers; capturing a real-world audio scene via one or more microphones to provide a second audio dataset for output via the one or more speakers; identifying, based on measured neural activity of the user, which of the first audio dataset and at least a portion of the real-world audio scene has the user's auditory attention; and based on the identification, controlling output of at least some of the first and / or second audio datasets via the one or more speakers.

[0018] In some example embodiments, the first audio data set may represent an audio track or a communication session received from a user device.

[0019] In some example embodiments, in response to identifying that the first audio data set has the user's auditory attention, controlling may include disabling or attenuating output of the second audio data set via the one or more speakers.

[0020] In some example embodiments, the method may further include generating a noise cancellation signal based on the captured real-world audio scene, wherein, further in response to identifying that the first audio data set has the user's auditory attention, the controlling includes enabling or increasing a gain associated with the noise cancellation signal for output to the one or more speakers.

[0021] In some example embodiments, the first audio data set may represent multiple audio sources; a first audio source of the multiple audio sources may be identified as having the user's auditory attention, and the first audio source may be amplified relative to the other(s) audio sources.

[0022] In some example embodiments, in response to identifying that at least a portion of the real-world audio scene has the user's auditory attention, the output may be controlled to disable or reduce a gain associated with the first audio data set.

[0023] In some example embodiments, in response to identifying that at least a portion of the real-world scene has the user's auditory attention, the output may be controlled to enable or increase a gain associated with at least some of the second audio data sets for output to the one or more speakers.

[0024] In some example embodiments, further in response to identifying that at least a portion of the real-world audio scene has the user's auditory attention, the output may be controlled to disable or reduce a gain associated with the noise cancellation signal for output to the one or more speakers.

[0025] In some example embodiments, the real-world audio scene may include multiple real-world audio sources, a first real-world audio source among the multiple real-world audio sources may be identified as having the user's auditory attention, and the method may further include: steering a sound capture beam of one or more microphones in the direction of the first real-world audio source so that an audio signal of the first real-world audio source is output to the one or more speakers with a higher gain than other(s) real-world audio sources.

[0026] In some example embodiments, the method may further include dividing the second audio data set into a plurality of frequency sub-bands, wherein a first frequency sub-band of the plurality of frequency sub-bands may be identified as having the user's auditory attention, and the output may be controlled to amplify the output of the first frequency sub-band with a higher gain than other frequency sub-bands.

[0027] In some example embodiments, further in response to identifying that the real-world audio scene has the user's auditory attention, the output may be controlled to disable or reduce a gain associated with a noise cancellation signal corresponding to a first frequency sub-band. In some example embodiments, the first frequency sub-band may correspond to speech audio.

[0028] In some example embodiments, the method may further include measuring a user's neural activity. In some example embodiments, the method may be performed by a headphone device.

[0029] In some example embodiments, the method may be performed using active noise cancellation functionality, the active noise cancellation functionality being operable in a transparent mode for outputting at least some of the second audio data set via one or more speakers, and controlling the output of at least some of the second audio data set may be configured to at least enable or control a gain at least associated with the transparent mode.

[0030] According to a third aspect, a computer program product is described comprising an instruction set which, when executed on an apparatus, is configured to cause the apparatus to perform a method comprising: outputting a first audio data set via one or more speakers; capturing a real-world audio scene via one or more microphones to provide a second audio data set for output via the one or more speakers; identifying, based on measured neural activity of the user, which of the first audio data set and at least a portion of the real-world audio scene has the user's auditory attention; and controlling output of at least some of the first and / or second audio data sets via the one or more speakers based on the identification.

[0031] In some example embodiments, the third aspect may include any other features mentioned with respect to the method of the second aspect.

[0032] According to a fourth aspect, an apparatus is described, comprising: at least one processing core, at least one memory including computer program code, the at least one memory and the computer program code being configured to, together with the at least one processing core, cause the apparatus to: output a first audio data set via one or more speakers; capture a real-world audio scene via one or more microphones to provide a second audio data set for output via the one or more speakers; identify, based on measured neural activity of the user, which of the first audio data set and at least a portion of the real-world audio scene has the user's auditory attention; and based on the identification, control output of at least some of the first and / or second audio data sets via the one or more speakers.

[0033] In some example embodiments, the fourth aspect may include any other features mentioned with respect to the method of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Example embodiments will be described by way of non-limiting examples with reference to the accompanying drawings, in which:

[0035] Figure 1 A system useful for understanding example embodiments is shown;

[0036] Figure 2A Shows that it can include Figure 1 a side view of a headphone device that is part of a system;

[0037] Figure 2B Shows the Figure 2A Headphones equipment;

[0038] Figure 3 shows a model for identifying a user's auditory attention according to some example embodiments;

[0039] Figure 4is a flowchart illustrating operations that may be performed according to some example embodiments;

[0040] Figure 5 showing a front view of a user listening to a first audio data set via a headphone device;

[0041] Figure 6 is a block diagram of a processing system according to some example embodiments;

[0042] Figure 7 is a block diagram of a noise cancellation system according to some example embodiments;

[0043] Figure 8A showing a front view of a user with a particular audio source of a real-world audio scene magnified;

[0044] Figure 8B Shown Figure 8A user, where microphone beamforming is used to amplify specific audio sources;

[0045] Figure 9 is a block diagram of an apparatus that may be configured according to one or more example embodiments; and

[0046] Figure 10 Non-transitory computer-readable media according to one or more example embodiments are shown. DETAILED DESCRIPTION

[0047] Disclosed herein are various example embodiments directed to controlling the output of audio data based on measured neural activity of a user.

[0048] Different audio data sets can represent different types of audio scenes. For example, a first audio data set can be output via one or more speakers and can represent a portion of an audio track or a communication session. The first audio data set can be stored in or streamed to a user device connected to one or more speakers, which can be speakers of a headphone device. For example, the second type of audio scene can include a real-world audio scene surrounding the user. The real-world audio scene can include one or more audio sources. At least a portion of the real-world audio scene can be captured via one or more microphones (e.g., one or more microphones of a headphone device) to provide a second audio data set.

[0049] As will be appreciated, some audio output devices (particularly headphone devices) provide active noise cancellation (ANC) systems. ANC systems can operate in two main modes: ANC mode and transparent (or pass-through) mode. In ANC mode, at least some of the audio data in the second audio data set can be processed to provide a noise cancellation signal, which can be output via one or more speakers to cancel or at least attenuate the surrounding real-world audio. The cancellation signal can be generated at least in part based on the second audio data set using a known ANC algorithm. As a result, the user will hear the first audio data set in an improved manner because less of the real-world audio will be heard. Generally speaking, the noise cancellation signal is primarily required to cancel lower audio frequencies because the physical structure of the headphone device can be able to acoustically block or attenuate higher audio frequencies. In transparent mode, the ANC mode can be disabled, and at least some of the audio data in the second audio data set can be output via one or more speakers. As a result, the user will hear at least some of the real-world audio around them, although they may still hear at least some of the first audio data set that may be output simultaneously. Transparency mode is sometimes used, for example, when a user wishes to hear and / or converse with another person nearby without having to pause playback of the first audio data set or remove the headphone device.

[0050] The first and second audio data sets can represent corresponding audio scenes, which can include one or more audio sources. An audio source can include any entity that emits audio (i.e., audible sound). Thus, an audio source can include a person, an animal, a musical instrument, a speaker, a vehicle, weather, or other environmental sounds. Such examples are not intended to be limiting.

[0051] Example embodiments may involve controlling the output of at least some of the second audio data set based on measured neural activity of a user. Example embodiments may, for example, involve identifying which of the first audio data set and at least a portion of a real-world audio scene has the user's auditory attention based on the user's measured neural activity. The identification may control the output of the at least some of the audio data in the second audio data set, for example according to the examples described below. Additionally or alternatively, example embodiments may involve controlling the output of at least some of the first audio data set based on the user's measured neural activity, for example by attenuating the output of the first audio data set if the at least a portion of the real-world audio scene identified as having the user's auditory attention.

[0052] Figure 1 is a block diagram of system 100 that may be useful in understanding example embodiments.

[0053] System 100 may include a server 110 , a user device 120 , a network 130 , and an audio output device, which in this example includes a headphone device 140 worn by a user 150 .

[0054] The headphone device 140 may include any form of head-mounted or ear-mounted audio output device, such as a pair of headphones, earbuds, a headset, or an extended reality (XR) headset. In this case, the audio data may be output to the left and right speakers using monaural, stereo, or (in the case of spatial audio data) binaural rendering.

[0055] Server 110 may be connected to user device 120 via network 130 to transmit audio data to the user device. Server 110 may, for example, comprise an Internet Protocol (IP) telecommunications server that transmits audio data comprising a portion of the speech of a communication session to user device 120. The communication session may, for example, comprise a voice call or conference call that may involve one or more participants in addition to user 150. The audio data may comprise spatial audio data encoded with spatial perception such that, when decoded and output to headphone device 140, one or more participants will be perceived in a generated audio scene at different respective positions relative to user 150. Example formats for spatial audio data may include, but are not limited to, multi-channel mixing, Ambisonics, parametric spatial audio (e.g., metadata-assisted spatial audio (MASA)), object-based audio, or any combination thereof. The spatial audio data may be encoded and decoded using a codec that may include, but is not limited to, the 3GPP Immersive Video and Audio Services (IVAS) format.

[0056] The transmission may be by means of any suitable streaming data protocol.

[0057] Alternatively or additionally, server 110 may provide one or more files comprising audio data to user device 120 for storage and processing at user device 120. The audio data may, for example, represent a soundtrack or audio associated with a video clip or movie.

[0058] At the user device 120, the audio data may be processed, rendered, and output to the headphone device 140. Alternatively, the audio data may be processed and rendered by the headphone device 140, such as where the headphone device 140 comprises part of an extended reality (XR) headset and thus has suitable processing capabilities.

[0059] In some example embodiments, user device 120 may include, but is not limited to, one of the following: a mobile phone, a tablet computer, a game console, a laptop computer, a personal computer, a vehicle navigation computer, or a wearable device. User device 120 may communicate with headset device 140 via a short-range communication channel (e.g., Bluetooth, Zigbee, WiFi, etc.). User device 120 may also include one or more cameras.

[0060] Network 130 may be any suitable data communications network, including, for example, one or more of: a radio access network (RAN), whereby communications with user equipment 120 are via one or more base stations; a WiFi network, whereby communications are via one or more access points; or a short-range network, such as a network using Bluetooth or Zigbee protocols.

[0061] Fig. 2 shows a first earphone 200 of the earphone device 140. It will be appreciated that a second earphone (not shown) of the earphone device may include the same or similar features.

[0062] The first earphone 200 may include a first portion 202 and a second portion 204 .

[0063] The first portion 202 may include a main body carrying a speaker 206 that, when in use, is located above, near, or partially within the ear canal of the user. The first portion 202 may carry one or more microphones 208 for capturing audio external to the user 150. For example, the audio captured by the microphone 208 may be processed as part of an active noise cancellation (ANC) function. Given the location of the microphone 208, the microphone 280 may be referred to as an external microphone, and the audio it captures may be used as part of a feedforward ANC algorithm. Although not shown, the first portion 202 may also carry one or more other "internal" microphones that are close to the speaker 206 and that may be used as part of a feedback ANC algorithm, but the feedforward case will be the focus of the description below.

[0064] The first portion 202 may also include a processing module 230 that may be configured to perform various operations described below. The processing module 230 may also provide the ANC functionality described above and may be configured to provide both ANC and transparent operating modes. Alternatively, a separate processing module may provide the ANC functionality.

[0065] The first portion 202 may also include one or more radio transceivers for communicating with the user equipment 120 .

[0066] The second portion 204 may include a curved arm that extends from the first portion 202 and, when in use, is positioned around the outside of the user's ear 220, such as Figure 2B shown.

[0067] The second portion 204 may include one or more sensors 210 that contact corresponding portions of the user's skin behind the ear 220. Alternatively or additionally, the first portion 202 may include one or more sensors that contact corresponding portions of the user's skin and / or are located within and in contact with the user's ear canal.

[0068] The one or more sensors 210 may sense or pick up biosignals (sometimes referred to as bioelectrical signals) generated by the user's brain during some form of activity. In an example embodiment, the activity may include listening to audio data as it is output via the speaker 206, or listening to a real-world audio scene while capturing at least a portion of the real-world audio scene.

[0069] Biosignals may include signals generated by a living being that can be measured and monitored. Such biosignals may, for example, represent changes in the current generated by the sum of potential differences across a specific tissue, organ, or cell system. For example, so-called electroencephalogram (EEG) signals may be picked up by one or more sensors 210 to provide an electrogram of the electrical activity of the user's brain in a non-invasive manner.

[0070] The biosignals picked up by the one or more sensors 210 may be provided to the processing module 230 .

[0071] Alternatively, the processing module 230 may be provided at the user device 120 , in which case the picked-up biosignal may be transmitted by the first earphone 200 to the user device.

[0072] The bio-signal may also be picked up by one or more sensors of a second earphone (not shown), and the bio-signal may be sent to the first earphone 200 (or the user device 120 ) including the processing module 230 .

[0073] The processing module 230 may include a machine learning (ML) model or be configured to access a machine learning model.

[0074] Figure 3 An example ML model 300 is shown. ML model 300 may be referred to as an attention model because it is configured to identify or estimate which audio (internal or external) has the user's current attention. ML model 300 may include, but is not limited to, an artificial neural network (ANN). There are various known types of ANNs, including feedforward neural networks, perceptron neural networks, convolutional neural networks, recurrent neural networks, deep neural networks, and the like. Each type may be more suitable for a specific application or task.

[0075] According to some example embodiments, the ML model 300 may be configured to receive the picked-up biosignal as a first input 302, which may be digitized and possibly pre-processed for measuring neural activity of the user 150 during the output and / or capture of at least a portion of the audio scene. The ML model 300 may also be configured to receive the aforementioned first audio dataset as a second input 303, which may represent at least a portion of an audio track or communication session being played via the speaker 206. The ML model 300 may also be configured to receive the aforementioned second audio dataset as a third input 304, which is captured by one or more microphones 208 during the output of the first audio dataset.

[0076] The ML model 300 can be configured to identify or classify which of the first audio data and the real-world audio (represented by the second audio dataset) has the auditory attention of the user 150 at a given time or within a given time window based on the received first, second, and third inputs 302, 303, 304. In other words, which of the internal or external real-world audio is currently being listened to by the user 150 based on the picked-up biosignals. For example, the ML model 300 can be configured according to the so-called auditory attention inference system described in Haghighi M, Moghadamfalahi M, Akcakaya M, Erdogmus D, “EEG-assisted Modulation of Sound Sources in the Auditory Scene” (Biomed Signal Process Control, January 2018; Vol. 39: 263-270). For example, the ML model 300 can determine or be informed of one or more audio signals from a corresponding internal audio source, one or more audio signals from a second audio data set, and one or more audio signals from a corresponding external (real-world) audio source in the first audio data set. For each of the internal and external audio sources, the ML model 300 can generate a corresponding probability value indicating the probability or likelihood that a particular audio source has the user's current attention. The corresponding probability value or the highest probability value can be provided as one or more output data sets 306.

[0077] For example, the ML model 300 can be trained using supervised or unsupervised learning methods. Supervised learning methods can involve measuring the user's biosignals while the user 150 is instructed, for example, to sequentially provide auditory attention to audio associated with each of a plurality of different audio sources and / or to audio sources having different respective locations relative to the user.

[0078] For example, another method for identifying that a particular audio source has the user's attention is to track the user's biosignals with corresponding audio signals for internal and external audio sources and / or cross-correlate the user's biosignals with corresponding audio signals for internal and external audio sources to identify the closest match.

[0079] Based on the above, it can be identified or at least estimated whether the user 150 is paying attention to (listening to) the first audio data set or at least a portion of the real-world scene represented by the second audio data set.

[0080] Figure 4 4 is a flow chart illustrating operation 400 according to one or more example embodiments. Operation 400 may be performed using hardware, software, firmware, or a combination thereof. For example, operation 400 may be performed individually or collectively by components, wherein the components may include at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the operation to be performed. Operation 400 may be performed, for example, by processing module 230 (which may be provided in headset device 140 or user device 120).

[0081] A first operation 401 may include outputting a first audio data set via one or more speakers.

[0082] The second operation 402 may include capturing a real-world audio scene external to the apparatus via one or more microphones to provide a second audio data set for output via one or more speakers.

[0083] A third operation 403 may include identifying which of the first audio data set and at least a portion of the real-world audio scene has the user's auditory attention based on the measured neural activity of the user.

[0084] A fourth operation 404 may include controlling output of at least some of the first and / or second audio data sets via one or more speakers based on the identification.

[0085] In some example embodiments, the first audio data set may represent an audio track or a communication session received from a user device.

[0086] Another operation may include measuring the user's neural activity. This may be performed using one or more ML models and / or using tracking and / or cross-correlation methods as described above.

[0087] The fourth operation 404 may include one or more different control options, which will now be described with reference to specific examples.

[0088] Reference to the second audio dataset as used herein may include a modified version of the second audio dataset, such as an amplified or filtered version of the second audio dataset.

[0089] Figure 5 A scenario is shown in which the user 150 listens to a first audio data set via a headphone device 502 .

[0090] The earphone device 502 may include a first earphone 502A and a second earphone 502B, each earphone being the same as or similar to the first earphone 200 of FIG. 2 , and at least one of the earphones may include the processing module 230 as described above.

[0091] The first audio data set may be received via a wireless link 522 (eg, a Bluetooth channel) from the user device 520. The first audio data set is decoded and output via the respective speakers of the first earphone 502A and the second earphone 502B.

[0092] Near the user 150 are a first (external) real-world audio source 504 and a second (external) real-world audio source 508 emitting respective first audio signals 506 and second audio signals 510 .

[0093] For example, the first audio source 504 may include a person, and the first audio signal 506 may represent speech. For example, the second audio source 508 may include an ambient sound source, and the second audio signal 510 may represent ambient sound. The first earphone 502A and the second earphone 502B may capture the first audio signal 506 and the second audio signal 510 via one or more corresponding microphones to generate a second audio data set.

[0094] Figure 6 Components of a first earphone 502A are shown according to an example embodiment. The first earphone 502A may include an antenna 602, one or more microphones 604, one or more speakers 606, one or more sensors 608, and a processing module 610.

[0095] Antenna 602 may be configured to communicate with Figure 5 The antenna 602 may, for example, comprise a Bluetooth antenna or the like. In other example embodiments, the user terminal 520 may communicate with the first headset 502A (and the second headset 502B) via a wired connection, in which case the antenna 602 may not be required.

[0096] The processing module 610 may include an attention model 620 (which may be configured to Figure 3 ), controller 625, and ANC system 630. ANC system 630 may operate in ANC mode and transparent mode.

[0097] In operation, a first audio data set may be received via antenna 602. The first audio data set may be provided as input to both the attention model 620 and the ANC system 630. The first audio signal 506 and the second audio signal 510 may be captured by one or more microphones 604, converted to a second audio data set via an analog-to-digital converter (ADC) (not shown), and provided as input to both the attention model 620 and the ANC system 630. One or more sensors 608 may pick up biosignals of the user 150 and provide these biosignals as input to the attention model 620.

[0098] User 150 may be listening to audio track 503, in other words, the first audio data set, at a first time. Attention model 620 may identify that the first audio data set has the auditory attention of user 150, for example, based on determining that the attention probability for the first audio data set is higher than the attention probability for the second audio data set. Attention model 620 may signal this identification via one or more sets of output data 633 (e.g., one or more probability values) that are input to controller 625.

[0099] The controller 625 can be configured to control the ANC system 630 using one or more control signals 640 according to the fourth operation 404. The controller 625 can, for example, output the one or more control signals 640 to cause the ANC system 630 to output at least some of the second audio data set via the one or more speakers 606. In this sense, controlling the output can include not outputting the second audio data set, outputting at least some of the second audio data set, or outputting the entirety of the second audio data set.

[0100] For example, in response to identifying that the first audio data set has the auditory attention of the user 150, the controller 625 can control the ANC system 630 to disable or attenuate the output of the second audio data set via the one or more speakers 606. One way to do this is to disable the transparency mode, or reduce the gain associated with the transparency mode to zero, or reduce the gain associated with the transparency mode toward zero. In some example embodiments, the controller 625 can substantially simultaneously control the ANC system 630 to enable or increase the gain associated with the noise cancellation signal for output to the one or more speakers 606, for example, by increasing the gain to a maximum level, or increasing the gain toward a maximum level.

[0101] At a later second time, the user 150 may attempt to listen to external audio, for example, a first audio signal 506 (e.g., speech) from a first audio source 504 (e.g., a person). The attention model 620 may, for example, identify that a portion of the real-world audio scene has the auditory attention of the user 150 based on determining that the attention probability for the second audio data set is higher than the attention probability for the first audio data set. The attention model 620 may signal this identification via one or more sets of output data 633 that are input to the controller 625.

[0102] For example, in response to identifying that a portion of the real-world audio scene has the auditory attention of user 150, controller 625 may control ANC system 630 to enable the gain associated with the transparency mode or increase the gain associated with the transparency mode to a maximum level, or increase the gain toward the maximum level. In some example embodiments, controller 625 may substantially simultaneously control ANC system 630 to disable or reduce the gain associated with the noise cancellation signal, for example, by reducing the gain to a zero level, or reducing the gain toward a zero level. Additionally or alternatively, for example, in an alternative embodiment where ANC system 630 is not provided, controller 625 may disable or reduce the gain associated with the first audio data set. In this way, the output of the first audio data set is attenuated so that user 150 hears the real-world audio scene.

[0103] Figure 7 An example form of ANC system 630 is shown, which may comprise an adapted form of a conventional ANC system.

[0104] The ANC system 630 may be configured to receive a second audio data set from the one or more microphones 604 , generate a noise cancellation signal via a first module 710 , and generate a so-called pass-through signal via a second module 720 , both of which are output via the one or more speakers 606 .

[0105] The pass-through signal may comprise the second audio data set or a modified version of the second audio data set, such as a filtered or post-processed version of the second audio data set. In this context, reference to the second audio data set as used herein may include such a modified version of the second audio data set.

[0106] Associated with each of the first module 710 and the second module 720 is a respective first mixer 712 and a second mixer 722 for controlling a first gain g1 and a second gain g2 associated with the noise cancellation signal and the pass-through signal, respectively. The controller 625 can be configured such that one or more control signals 640 control the respective first gain g1 and second gain g2. For example, the first gain g1 and second gain g2 can be controlled in opposite directions such that one gain increases as the other decreases.

[0107] For example, if the attention model 620 identifies that the first audio data set has the auditory attention of the user 150, the first gain g1 can be set to a maximum level (e.g., 1) or tends to a maximum level (e.g., 1), and the second gain g2 can be set to zero or other minimum value or tends to zero or other minimum value.

[0108] For example, if the attention model 620 identifies that a portion of the real-world audio scene has the auditory attention of the user 150, the first gain g1 can be set to zero or other minimum value or tends to zero or other minimum value, and the second gain g2 can be set to a maximum level (e.g., 1) or tends to a maximum level (e.g., 1).

[0109] In some example embodiments, the first audio data set may represent a plurality of audio sources.

[0110] refer to Figure 8A For example, a first audio dataset may represent a first participant 810 and a second participant 820 of a communication session, which may be output at different respective locations relative to the user 150. The first participant 810 and the second participant 820 may be represented in the first audio dataset as respective spatial audio objects that the attention model 620 may distinguish.

[0111] If attention model 620 identifies that first participant 810 has the auditory attention of user 150, then in addition to the operations described above for the case where the first audio data set has the user's auditory attention, controller 625 may also cause the gain associated with the audio of first participant 810 to be increased so that it will be louder than the audio of second participant 820, as indicated by dotted line 830. Controller 625 may communicate with user device 520 for this purpose, or may control the gain itself.

[0112] In some example embodiments, the second audio dataset represents a plurality of external real-world audio sources, e.g. Figure 5 , the attention model 620 can identify that a particular real-world audio source has the auditory attention of the user 150. This can be achieved if the one or more external microphones include a microphone array that enables audio signals from different directions to be captured and distinguished by the attention model 620.

[0113] Figure 8BA situation is shown in which a first audio source (person) 504 has the auditory attention of the user 150, in which case, in addition to the operations described above for the situation in which the second audio data set has the user's auditory attention, the controller 625 can also cause one or more microphones to steer the sound capture beam 902 in the direction of the first audio source 504 so that the audio signal of the first audio source 504 will be amplified with a higher gain than the second audio source 508 or indeed any other external audio source.

[0114] In some example embodiments, the second audio data set may be divided into a plurality of frequency sub-bands.

[0115] The attention model 620 can identify that a first frequency subband (e.g., corresponding to speech audio) among a plurality of frequency subbands has the auditory attention of the user 150. For example, the attention model 620 can identify that the user 150 is listening to speech, which it knows corresponds to a frequency subband of approximately 400Hz-4kHz or 400Hz-8kHz (for higher speech quality). In addition to the above-mentioned operations for the case where the second audio data set has the auditory attention of the user 150, the controller 625 can also amplify the output of the first frequency subband with a higher gain than other frequency subbands. This can be performed by bandpass filtering the passthrough signal using the first frequency subband. In some example embodiments, the controller 625 can also disable or reduce the gain associated with the noise cancellation signal corresponding to the first frequency subband. This can involve bandstop filtering the noise cancellation signal using the first frequency subband. The controller 625 can communicate with the user device 520 for this purpose, or can control the gain internally.

[0116] Although example embodiments are described with respect to controlling the ANC system 630, other example embodiments may be directed solely to controlling whether the second audio data set is output via one or more speakers.

[0117] Thus, example embodiments provide an improved automated way to emphasize the output of audio signals based on which audio data set the user 150 is listening to. The user 150 does not need to perform direct or conscious interaction. Example embodiments may also provide energy savings by only amplifying or passing through the audio data that the user 150 is listening to.

[0118] Example device

[0119] Figure 9An example apparatus 900 capable of supporting at least some embodiments is shown. Device 900 is shown and may include a headset device, such as first headset 200 or user device 120. Device 900 includes a processor 910, which may include, for example, a single-core or multi-core processor, where a single-core processor includes one processing core and a multi-core processor includes more than one processing core. Processor 910 may generally include a control device. Processor 910 may include more than one processor. Processor 910 may be a control device. A processing core may include, for example, a Cortex-A8 processing core manufactured by ARM Holdings or a Steamroller processing core manufactured by Advanced Micro Devices. Processor 910 may include at least one Qualcomm Snapdragon and / or Intel Atom processor. Processor 910 may include at least one application-specific integrated circuit (ASIC). Processor 910 may include at least one field-programmable gate array (FPGA). Processor 910 may be a component for executing method steps in device 900. Processor 910 may be configured, at least in part, by computer instructions to perform actions.

[0120] The processor may include circuitry, or be constructed as one or more circuitry, configured to perform the various stages of the method according to the embodiments described herein. As used in this application, the term "circuitry" may refer to one or more or all of the following: (a) a hardware circuit implementation only, for example, an implementation using only analog and / or digital circuitry, and (b) a combination of hardware circuitry and software, for example, as applicable: (i) a combination of (one or more) analog and / or digital hardware circuits with software / firmware, and (ii) any portion of (one or more) hardware processors (including (one or more) digital signal processors), software, and (one or more) memories with software, which work together to enable a device or apparatus configured to control its function to perform various functions, and (c) (one or more) hardware circuits and / or (one or more) processors, such as (one or more) microprocessors or portions of microprocessors, which require software (e.g., firmware) to operate, but the software may not be present when not required for operation.

[0121] This definition of "circuitry" applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term "circuitry" also covers an implementation of merely a hardware circuit or processor (or multiple processors), or a portion of a hardware circuit or processor, and its accompanying software and / or firmware. The term "circuitry" also covers (for example, and if applicable to a particular claim element) a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or networking equipment.

[0122] Device 900 may include memory 920. Memory 920 may include random access memory and / or permanent memory. Memory 920 may include at least one RAM chip. For example, memory 920 may include solid-state, magnetic, optical, and / or holographic memory. Memory 920 may be at least partially accessible to processor 910. Memory 920 may be at least partially included in processor 910. Memory 920 may be a component for storing information. Memory 920 may include computer instructions that processor 910 is configured to execute. When computer instructions configured to cause processor 910 to perform certain actions are stored in memory 920, and device 900 is generally configured to operate under the direction of processor 910 using computer instructions from memory 920, processor 910 and / or at least one of its processing cores may be considered to be configured to perform the certain actions. Memory 920 may be at least partially included in processor 910. Memory 920 may be at least partially external to device 900, but accessible to device 900.

[0123] The device 900 may include a transmitter 930. The device 900 may include a receiver 940. The transmitter 930 and the receiver 940 may be configured to transmit and receive information, respectively, according to at least one cellular or non-cellular standard.

[0124] Transmitter 930 may include more than one transmitter. Receiver 940 may include more than one receiver. For example, transmitter 930 and / or receiver 940 may be configured to operate in accordance with Global System for Mobile Communications (GSM), Wideband Code Division Multiple Access (WCDMA), 5G / NR, 5G-Advanced (i.e., NR Rel-18, 19 and higher), Long Term Evolution (LTE), IS-95, Wireless Local Area Network (WLAN), Ethernet, and / or Worldwide Interoperability for Microwave Access (WiMAX) standards.

[0125] The device 900 may include a near field communication (NFC) transceiver 950. The NFC transceiver 950 may support at least one NFC technology, such as NFC, Bluetooth, Wibree, or similar technology.

[0126] The device 900 may include a user interface UI 960. The UI 960 may include at least one of a display, a keyboard, a touch screen, a vibrator arranged to signal the user by vibrating the device 900, a speaker, and a microphone. The user may be able to operate the device 900 via the UI 960, for example, to accept an incoming phone call, make a phone call or video call, browse the internet, manage digital files stored in the memory 920 or on a cloud accessible via the transmitter 930 and receiver 940 or via the NFC transceiver 950, and / or play games.

[0127] Device 900 may include or be arranged to accept a subscriber identification module 970. Subscriber identification module 970 may include, for example, a subscriber identity module (SIM) card that may be installed in device 900. Subscriber identification module 970 may include information identifying a subscription of a user of device 900. Subscriber identification module 970 may include encryption information that may be used to verify the identity of the user of device 900 and / or facilitate encryption of transmitted information and billing of the user of device 900 for communications accomplished via device 900.

[0128] The processor 910 may be equipped with a transmitter that is arranged to output information from the processor 910 to other devices included in the device 900 via wires within the device 900. Such a transmitter may include a serial bus transmitter that is arranged to output information to the memory 920 for storage therein, for example, via at least one wire. As an alternative to a serial bus, the transmitter may include a parallel bus transmitter.

[0129] Likewise, the processor 910 may include a receiver arranged to receive information in the processor 910 from other devices included in the device 900 via wires internal to the device 900. Such a receiver may include a serial bus receiver arranged to receive information from the receiver 940, for example, via at least one wire, for processing in the processor 910. As an alternative to a serial bus, the receiver may include a parallel bus receiver.

[0130] The device 900 may also include Figure 9 Other devices not shown in the figure. For example, if device 900 comprises a smartphone, it may include at least one digital camera. Some devices 900 may include a rear-facing camera and a front-facing camera, wherein the rear-facing camera may be used for digital photography and the front-facing camera is used for video calling. Device 900 may include a fingerprint sensor, which is arranged to at least partially authenticate the user of device 900. In some embodiments, device 900 lacks at least one of the components described above. For example, some devices 900 may lack NFC transceiver 950 and / or user identification module 970.

[0131] The processor 910, memory 920, transmitter 930, receiver 940, NFC transceiver 950, UI 960, and / or user identification module 970 can be interconnected in a variety of different ways via wires within the device 900. For example, each of the above components can be connected to a main bus within the device 900 to allow these components to exchange information. However, as will be appreciated by those skilled in the art, this is merely an example, and depending on the embodiment, different ways of interconnecting the at least two components described above can be selected without departing from the scope of the present invention.

[0132] Figure 10 A non-transitory medium 1000 according to some embodiments is shown. Non-transitory medium 1000 is a computer-readable storage medium. It may be, for example, a CD, a DVD, a USB drive, a Blu-ray disc, etc. Non-transitory medium 1000 stores computer program instructions that cause a device to perform any of the aforementioned methods, such as those disclosed with respect to the flowcharts and related features in this specification.

[0133] In one or more embodiments, the described features, structures or characteristics may be combined in any suitable manner. In the previous description, many specific details (such as examples of length, width, shape, etc.) are provided to provide a thorough understanding of embodiments of the present invention. However, those skilled in the relevant art will recognize that the present invention may be practiced without one or more specific details or using other methods, components, materials, etc. In other cases, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the present invention.

[0134] Although the above examples illustrate the principles of the embodiments in one or more specific applications, it will be apparent to those skilled in the art that many modifications can be made in the form, use, and details of implementation without inventiveness and without departing from the principles and concepts of the invention. Therefore, the present invention is not limited except by the claims set forth below.

[0135] The verbs "to comprise" and "to include" are used in this document as open limitations that neither exclude nor require the presence of unrecited features. Features described in the dependent claims may be freely combined with each other unless expressly stated otherwise. Furthermore, it will be understood that the use of "a" (i.e., the singular) in this document does not exclude a plurality.

Claims

1. A device comprising: means for outputting a first audio data set via one or more speakers of the apparatus; means for capturing a real-world audio scene external to the apparatus via one or more microphones of the apparatus to provide a second audio data set for output via the one or more speakers; means for identifying which of the first audio data set and at least a portion of the real-world audio scene has the user's auditory attention based on the user's measured neural activity; means for generating a noise cancellation signal based on the captured real-world audio scene; means for controlling output of at least some of the first audio data set and / or the second audio data set via the one or more speakers based on the identifying; as well as In response to identifying that the first audio data set has the user's auditory attention, the means for controlling output is configured to enable or increase a gain associated with the noise cancellation signal for output to the one or more speakers.

2. The device according to claim 1, wherein The first audio data set represents an audio track or a communication session received from a user device associated with the apparatus.

3. The device according to claim 1, wherein The first audio dataset represents a plurality of audio sources; The means for identifying is configured to: identify that a first audio source among the plurality of audio sources has the auditory attention of the user; as well as The means for controlling the output is configured to amplify the first audio source relative to other audio sources.

4. An apparatus according to any preceding claim, wherein In response to identifying that at least a portion of the real-world audio scene has the user's auditory attention, the means for controlling the output is configured to disable or reduce a gain associated with the first audio data set.

5. An apparatus according to any preceding claim, wherein In response to identifying that at least a portion of the real-world scene has the user's auditory attention, the component for controlling the output is configured to: enable or increase a gain associated with at least some of the second audio data set for output to the one or more speakers.

6. The device according to claim 5, wherein The means for controlling output is configured to disable or reduce a gain associated with the noise cancellation signal for output to the one or more speakers further in response to identifying that at least a portion of the real-world audio scene has the user's auditory attention.

7. The device according to claim 5 or 6, wherein: The real-world audio scene includes a plurality of real-world audio sources, The means for identifying is configured to: identify that a first real-world audio source of the plurality of real-world audio sources has the auditory attention of the user, and The apparatus further includes means for steering sound capture beams of the one or more microphones in the direction of the first real-world audio source so that the audio signal of the first real-world audio source is output to the one or more speakers with a higher gain than other real-world audio sources.

8. The apparatus according to any one of claims 5 to 7, further comprising: means for dividing the second audio data set into a plurality of frequency sub-bands, wherein The means for identifying is configured to: identify a first frequency sub-band of the plurality of frequency sub-bands as having the user's auditory attention, and The means for controlling the output is configured to amplify the output of the first frequency sub-band with a higher gain than other frequency sub-bands.

9. The device according to claim 8, wherein The means for controlling an output is configured to, further in response to identifying that the real-world audio scene has the user's auditory attention, disable or reduce a gain associated with a noise cancellation signal corresponding to the first frequency subband.

10. The device according to claim 9, wherein The first frequency sub-band corresponds to speech audio.

11. The apparatus according to any preceding claim, further comprising: Means for measuring neural activity of the user.

12. An apparatus according to any preceding claim, wherein The arrangement is comprised by a headphone device.

13. The device according to claim 12, wherein the apparatus comprising active noise cancellation functionality operable in a transparent mode for outputting said at least some of the second audio data set via the one or more speakers, and The means for controlling the output of the at least some of the second audio data sets is configured to at least enable or control a gain associated with at least the transparency mode.

14. A method comprising: outputting a first audio data set via one or more speakers; capturing a real-world audio scene via one or more microphones to provide a second audio data set for output via the one or more speakers; identifying, based on the user's measured neural activity, which of the first audio data set and at least a portion of the real-world audio scene has the user's auditory attention; generating a noise cancellation signal based on the captured real-world audio scene; controlling output of at least some of the first audio data set and / or the second audio data set via the one or more speakers based on the identifying; as well as In response to identifying that the first audio data set has the user's auditory attention, controlling includes enabling or increasing a gain associated with the noise cancellation signal for output to the one or more speakers.

15. A computer program product comprising a set of instructions, the set of instructions being configured, when executed on a device, to cause the device to perform a method comprising: outputting a first audio data set via one or more speakers; capturing a real-world audio scene via one or more microphones to provide a second audio data set for output via the one or more speakers; identifying, based on the user's measured neural activity, which of the first audio data set and at least a portion of the real-world audio scene has the user's auditory attention; generating a noise cancellation signal based on the captured real-world audio scene; controlling output of at least some of the first audio data set and / or the second audio data set via the one or more speakers based on the identifying; as well as In response to identifying that the first audio data set has the user's auditory attention, controlling includes enabling or increasing a gain associated with the noise cancellation signal for output to the one or more speakers.