Context-aware adaptive control of sound leakage

A context-aware system dynamically adjusts audio characteristics based on detected volume and user context to minimize sound leakage, addressing disturbances in multi-user environments by quantitatively managing audio output.

WO2025149400A1PCT designated stage expired Publication Date: 2025-07-17NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/050013
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2025-01-02
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Sound leakage occurs when audio escapes from enclosed areas, causing disturbance to unintended listeners due to increased volume, particularly in environments with thin or poorly insulated walls, where individuals are subjected to distorted and disruptive sounds.

Method used

A context-aware system that detects audio volume and user context to adjust audio characteristics dynamically, using devices to reduce sound leakage by determining volume thresholds and generating signals for volume adjustments or notifications.

Benefits of technology

Effectively minimizes sound leakage by quantitatively adapting audio output based on the context of secondary users, reducing disturbance and enhancing audio regulation in multi-user environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025050013_17072025_PF_FP_ABST
    Figure EP2025050013_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein is an apparatus, comprising: means for obtaining volume data indicative of a volume of audio detected at a computing device remote from the apparatus; means for obtaining context information indicative of a context of a user associated with the computing device; means for determining, based on the received context information, a volume threshold; and means for generating a signal when the volume of the detected audio exceeds the determined volume threshold. A method is also disclosed, comprising: detecting audio, wherein the audio is output by an audio outputting apparatus to a primary user; obtaining a volume of the detected audio; obtaining a context of a secondary user; determining, based on the context information, a volume threshold; generating a signal when the volume of the detected audio exceeds the determined volume threshold.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Context-Aware Adaptive Control of Sound Leakage

[0002] Field

[0003] Example embodiments may relate to apparatus, systems, methods and / or computer programs for performing context-aware adaptive control of sound leakage.

[0004] Background

[0005] Sound leakage occurs when sound escapes out of an enclosed area / room. This occurs in headphones / earphones / earbuds hardware and speaker setups, where sound leaks out of the ear or through the walls of a room in which an audio outputting apparatus is located. As the volume of the sound increases, the leakage increases.

[0006] Sound leakage is highly pertinent in homes and other buildings with thin or poorly insulated walls, where people in the vicinity of audio playback who are not the intended listeners can therefore be subject to distorted and / or disturbing sounds. An example is a parent taking a meeting from a home office while their children are playing video games, whilst noise from the video games also leaks through the walls and distracts the parent. Similarly, this could occur with roommates or housemates playing movies / listening to music late at night while others are trying to sleep.

[0007] Summary

[0008] The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention.

[0009] According to a first aspect, there is described an apparatus, comprising: means for obtaining volume data indicative of a volume of audio detected at a device; means for obtaining context information indicative of a context of a user associated with the device; means for determining, based on the context information, a volume threshold; and means for generating a signal when the volume of the detected audio exceeds the determined volume threshold. In some examples, the device is remote from the apparatus. In some examples, the apparatus further comprises means for outputting the audio with one or more audio characteristics. In some examples, the apparatus further comprises means for adjusting the one or more audio characteristics, based on the generated signal, to reduce the volume of audio detectable at the device. In some examples, the means for adjusting the one or more audio characteristics comprises means for adjusting the audio characteristics based on historical volume data associated with the device. In some examples, the apparatus further comprises means for providing a user notification based on the generated signal, the notification indicating that the volume of audio detected at the device exceeds the volume threshold.

[0010] In some examples, the apparatus further comprises means for providing, to another device configured to output audio with one or more audio characteristics, the generated signal. Optionally, the generated signal comprises one or more commands which cause the other device to: adjust the one or more audio characteristics to reduce the volume of audio detectable at the device; and / or provide a user notification, the notification indicating that the volume of audio detected at the device exceeds the volume threshold.

[0011] In some examples, the one or more audio characteristics comprise one or more of: a volume, a bass level, a middle level, a treble level, or an equalization. Adjusting a bass level can comprise adjusting a volume of the audio within the bass frequency range of 60 to 250 Hz. Adjusting a middle level can comprise adjusting a volume of the audio within the midrange frequency range of 500 to 2000 Hz. Adjusting a treble level can comprise adjusting a volume of the audio within the presence frequency range of 4 to 6 kHz.

[0012] In some examples, the apparatus is comprised by the device, the apparatus comprising means for detecting the audio. Optionally, the audio is output by another device remote from the apparatus, the apparatus further comprising means for receiving an indication of the audio being output by the other device, wherein the means for obtaining the volume data comprises means for obtaining the volume data based on the received indication. In some examples, the apparatus further comprises means for transmitting the generated signal.

[0013] In some examples, the context comprises a current activity of the user and / or an expected future activity of the user. Optionally, the activity comprises: sleeping, consuming media, or attending a meeting.

[0014] In some examples, the generated signal comprises the volume data and the volume threshold. In some examples, the generated signal comprises an indication of a difference between the detected volume and the volume threshold. In some examples, the generated signal comprises a control signal indicative of an adjustment of the audio, the adjustment determined based on the detected volume and the volume threshold. According to a second aspect, there is provided a system, comprising: an apparatus comprising means for outputting audio with one or more audio characteristics; a device comprising means for detecting the output audio; means for obtaining volume data indicative of a volume of the detected audio; means for obtaining context information indicative of a context of a user associated with the device; means for determining, based on the received context information, a volume threshold; and means for generating a signal when the volume of the detected audio exceeds the determined volume threshold.

[0015] Optionally, the apparatus further comprises means for adjusting the one or more audio characteristics, based on the generated signal, to reduce the volume of audio detectable at the device. Optionally, the system further comprises means for providing a user notification based on the generated signal, the notification indicating that the volume of audio detected at the device exceeds the volume threshold. In some examples, the device is a smart speaker. In some examples, the apparatus is a smart speaker.

[0016] Also described is a method comprising: obtaining volume data indicative of a volume of audio detected at a device; obtaining context information indicative of a context of a user associated with the device; determining, based on the context information, a volume threshold; and generating a signal when the volume of the detected audio exceeds the determined volume threshold.

[0017] Also disclosed is a computer program comprising instructions for causing an apparatus to perform at least the steps of: obtaining volume data indicative of a volume of audio detected at a device; obtaining context information indicative of a context of a user associated with the device; determining, based on the context information, a volume threshold; and generating a signal when the volume of the detected audio exceeds the determined volume threshold.

[0018] Also disclosed is a computer-readable medium (such as a non-transitory computer- readable medium) comprising program instructions stored thereon for performing at least the following: obtaining volume data indicative of a volume of audio detected at a device; obtaining context information indicative of a context of a user associated with the device; determining, based on the context information, a volume threshold; and generating a signal when the volume of the detected audio exceeds the determined volume threshold.

[0019] Also disclosed is an apparatus comprising: at least one processor; and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to: obtain volume data indicative of a volume of audio detected at a device; obtain context information indicative of a context of a user associated with the device; determine, based on the context information, a volume threshold; and generate a signal when the volume of the detected audio exceeds the determined volume threshold.

[0020] Also disclosed is an apparatus comprising: a first obtaining module configured to obtain volume data indicative of a volume of audio detected at a device; a second obtaining module configured to obtain context information indicative of a context of a user associated with the device; a first determining module configured to determine, based on the context information, a volume threshold; and a first generating module configured to generate a signal when the volume of the detected audio exceeds the determined volume threshold.

[0021] Also disclosed is a system comprising an apparatus and a device. The apparatus comprises a first outputting module configured to output audio with one or more audio characteristics. The device comprises: a first detecting module configured to detect the output audio; a first obtaining module configured to obtain volume data indicative of a volume of the detected audio; a second obtaining module configured to obtain context information indicative of a context of a user associated with the device; a first determining module configured to determine, based on the received context information, a volume threshold; and a first generating module configured to generate a signal when the volume of the detected audio exceeds the determined volume threshold.

[0022] Also disclosed is a system comprising an apparatus and a device. The apparatus comprises at least one processor and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to output audio with one or more audio characteristics. The device comprises at least one processor and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to: detect the output audio; obtain volume data indicative of a volume of the detected audio; obtain context information indicative of a context of a user associated with the device; determine, based on the received context information, a volume threshold; and generate a signal when the volume of the detected audio exceeds the determined volume threshold.

[0023] The embodiments or examples of each of the above aspects can be combined with any of the other aspects in any suitable combination. Brief description of the drawings

[0024] Example embodiments will now be described, by way of example only, with reference to the following schematic drawings, in which:

[0025] FIG. 1A is a schematic block diagram of an example system comprising an audio outputting apparatus , with a primary user listening to the audio, and one or more devices associated with a secondary user who is not actively listening to the audio;

[0026] FIG. IB is a schematic block diagram of another example system comprising an audio outputting apparatus, with a primary user listening to the audio, and a one or more devices associated with a secondary user who is not actively listening to the audio;

[0027] FIG. 10 is a schematic block diagram of another example system comprising an audio outputting apparatus, with a primary user listening to the audio, and two devices associated with a secondary user who is not actively listening to the audio;

[0028] FIG. ID is a schematic block diagram of another example system comprising an audio outputting apparatus, with a primary user listening to the audio, and one or more device(s) associated with a secondary user who is not actively listening to the audio; FIG. 2 is an example flow chart;

[0029] FIG. 3 illustrates an example implementation of operations of the flow chart of FIG. 2.

[0030] FIG. 4 is a schematic view of an apparatus which may be configured according to one or more example implementations of the process described herein; and

[0031] FIG. 5 is a plan view of non-transitory media. In the description and drawings, like reference numerals refer to like elements throughout.

[0032] Detailed description The following description relates to an adaptive approach to controlling sound leakage, wherein a volume of an active audio outputting apparatus is controlled in a context aware manner.

[0033] The volume of an audio output can be adapted or adjusted based on the location / distance of a primary listener from an audio outputting apparatus or based on ambient environmental noise. However, in environments where there are multiple users, the primary listener might not know if other users are present, what those users are doing, or how much of the audio output they can hear. It can therefore be difficult for the primary listener of the audio playing device to gauge how loud they can play audio content without disturbing others.

[0034] As described herein, an apparatus is provided which comprises means for obtaining volume data indicative of a volume of audio detected at a device, means for obtaining context information indicative of a context of a user associated with the device, means for determining, based on the context information, a volume threshold, and means for generating a signal when the volume of the detected audio exceeds the determined volume threshold. The signal can be an internal control signal, or an external signal which causes one or more actions to be performed at one or more other devices. By providing a context-aware approach to adapting volume, which takes into account the volume heard by other users and a context associated with those users, sound or audio leakage can be dynamically controlled. This approach recognizes that sound leakage can be, but is not always, a problem for other users. Moreover, by using the volume detected at a location remote from the audio outputting apparatus to determine the volume adjustments, the adjustments are based on a quantitative measure and not a subjective assessment of the primary listener. The process described herein can be repeated over time, allowing for dynamic, adaptive control of sound-leakage. In this way, the adverse effects of sound leakage can be reduced or minimized.

[0035] The present approach thus leverages user devices to (optionally automatically) adjust audio characteristics based on the context of other users located near the audio source. Situations where one person plays music or watches television, but others are subject to sound leakage from the audio are common. Approaches for volume regulation are known. However, the present approach recognises that the identification of the relevant contexts can indicate what amount of audio sound leakage is acceptable for users, which can improve the regulation of the volume of sound leakage detected by the users, allowing to dynamically counteract sound leakage.

[0036] With reference to FIG. 1A, a first example of context-aware adaptive control of sound leakage will be described. The first example includes a system 400 comprising an audio outputting apparatus 100a and a secondary device 200, as described below.

[0037] The audio outputting apparatus 100a is also called herein a "master" device, since it is the device at which the audio playback can be controlled.

[0038] The apparatus 100a is configured to output audio 300 with one or more audio characteristics. A primary user (active listener) 310 is located proximate to, or in an environment of, the apparatus 100a and is actively listening to the audio 300. The apparatus 100a can be e.g., a television, radio, smart speaker, smart phone, tablet, mobile computing device, or any other device configured with means 112a for outputting the audio with the one or more audio characteristics. The one or more audio characteristics can comprise one or more of: a volume, a bass level, a middle level, a treble level, or an equalization. A bass level is the volume of the audio within the bass frequency range of 60 to 250 Hz. A middle level is the volume of the audio within the midrange frequency range of 500 to 2000 Hz. A treble level is the volume of the audio within the presence frequency range of 4 to 6 kHz. The audio characteristic(s) characterise the sound / audio which is output by the means 112a of the apparatus 100a.

[0039] In this example, device 200 is located remote from the apparatus 100a. Although shown here as a single device, device 200 can be implemented as one or more devices, as required. The device might be another part of a same room / environment as the apparatus 100a, or in another room. A secondary user (non-active listener) 312 is associated with the device 200. In some situations, the secondary user 312 can hear parts of the audio 300 that the primary user is listening to, which may be undesirable. Although the primary user 310 can adjust the audio characteristics to minimise inconvenience to the secondary user 312, the primary user is not aware of how loud the audio is at the location of the secondary user 312. The device 200 is provided with means for detecting 202 the audio 300 which is output by the apparatus 100a. The device 200 can be any audio sensor or computing device suitable for detecting the audio 300. The device 200 can be e.g., a television, radio, smart speaker, smart phone, tablet, mobile computing device, an audio sensor, a microphone on a wire, or any other device with means 202 to detect the audio. The means 202 for detecting the audio can comprise a microphone (stand alone or implemented into a computing device), or any other suitable sensor (such as an audio transducer). In this specific example of FIG. 1A, the device is implemented as a computing device 200 which is remote from the audio outputting apparatus 100a. In addition to the detecting means 202, the device 200 may comprise one or more additional modules or means, as described below. These different modules or means can be implemented in any suitable combination, depending on the functionality of the other devices and apparatus of system 400. For example, when functionality described with reference to the device 200 is implemented on the apparatus 100a, the corresponding modules may not be present on the device 200, or may be present but not used. Similarly, when functionality described with reference to the apparatus is implemented on the device, the corresponding modules may not be present on the apparatus, or may be present but not used.

[0040] The computing device 200 of this example further comprises means for determining 204 a volume of the detected audio. By determining the volume at the computing device, an impact of the audio on the secondary user 312 can be better understood.

[0041] Optionally, the computing device 200 further comprises means for receiving 212 an indication 308 of the audio being output by the remote apparatus. The means for determining a volume of the detected audio can then determine the volume of the detected audio based on the received indication. The combination of means 212, 204 can help the device 200 to specifically determine which components of the ambient detected audio are audio which has leaked from the audio playback apparatus 100a, which can help to improve a determination of the volume of the detected audio. This indication can be received directly from the master device / audio outputting apparatus 100a, or from another device (such as a home hub or entertainment centre associated with the apparatus 100a).

[0042] In this example, the computing device further comprises means for transmitting 206 volume data 302, where the volume data is indicative of the volume of the detected audio. In this example, means 206 are configured to transmit the volume data 302 directly to the apparatus 100a, but it will be understood that the volume data may be transmitted between the computing device and the apparatus via one or more intermediate devices, as appropriate. As discussed, the volume data 302 is indicative of a volume of the audio 300 detected at the computing device 200, which is remote from the apparatus 100a.

[0043] The apparatus 100a comprises means for obtaining 102 the volume data. The volume data is in this example received from the computing device 200, as shown in FIG. 1A. Alternatively, in other examples not illustrated here, the volume data is received at the apparatus 100a from one or more other devices; these could be other devices that have actively detected the volume of the audio (i.e., separately from the computing device) or devices that act as intermediaries between the computing device 200 and the apparatus 100a.

[0044] In still other examples not illustrated here, the means for obtaining 102 the volume data can comprise the functionality of means 204 described above. For example, the device 200 can detect the audio via means 202, but the volume data 302 of the detected audio is determined directly by the apparatus 100a by the means 102 for obtaining the volume data. It will be understood that in such an arrangement, information indicative of the detected audio can be received by the apparatus 100a (directly or indirectly) from the device 200 to facilitate the volume data being determined by the apparatus 100a. In such an arrangement, the device 200 may not comprise means 204. Device 200 also may not comprise means 206, or means 206 may be configured instead to transmit either the detected audio or information indicative of the detected audio so as to allow the means 102 to obtain the volume data 302.

[0045] By obtaining the volume data, the apparatus 100a and / or primary user 310 can better understand how the audio 300 impacts the secondary user. However, the actual impact of the sound leakage from audio 300 on the secondary user is dependent in part on what the secondary user is actually doing. In particular, certain activities that the secondary user 312 is performing are more likely to be disturbed, disrupted or interfered with than others.

[0046] Therefore, the apparatus further comprises means for obtaining 104 context information 304 indicative of a context of the secondary user 312 associated with the computing device 200. The context information is used to provide context aware adaptation of the volume of the audio 300, and either indicates or is used to determine an appropriate volume level for that user. The context information can be received from any suitable device, including but not limited to computing device 200. In some examples (not shown here), the context information is received at the apparatus from e.g., a remote cloud processing device or other device which is connected to or associated with the apparatus 100a (e.g., the devices can be linked to one another, or to a common device such as a Wi-Fi router). In some examples, the apparatus 100a is a smart speaker and it receives the volume and / or context information from one or more other smart speakers located in different rooms or different parts of the environment.

[0047] In some examples, when the apparatus 100a starts to play audio 300, it can actively search for nearby devices to detect the presence of users nearby or can determine nearby devices from stored information. This can be considered a "discovery" phase. For example, apparatus can search for users' smartphones or wearables, or users' detected proximity to static devices such as smart speakers. The context of each user is received or determined by the apparatus 100a, e.g., based on sensors on their smartphone / wearable, and / or sensors of other devices proximate to the user. For example, the sensors embedded in mobile devices can be used to detect contextual factors associated with respective (secondary) users. In some examples, the apparatus 100a notifies the detected nearby devices (such as computing device 200 or another intermediary device, for example) to run the context recognition process. This context recognition process can be run periodically, as discussed below. If no devices owned by other people are found nearby, then no adjustments to the audio characteristics are needed. On the other hand, if devices are found, certain user contexts might instead require adjustment to regulate the volume of leaked audio. Audio adaptation based on context may be useful in e.g., a home office scenario, a shared home where roommates have different waking and sleeping hours, university students sharing a home and needing to study for exams, etc. Regardless of the specific scenario, the sensors in computing devices such as personal computing devices and / or wearables can be activated to determine these contexts. The context determination can happen before, after or simultaneously with the volume detection discussed above. In some examples, the context is stored and then accessed when the volume regulation process begins.

[0048] In other examples, the determination of user context can be initiated in response to detecting audio at device 200, or in response to user initiation of the volume regulation process. In these examples, the "discovery" phase may not be initiated by apparatus 100a, but by the individual computing device 200. For example, computing device 200 may initiate the volume adaptation process by providing volume data and / or context information to the apparatus 100a. In the specific example shown in FIG. 1A, the computing device 200 is configured with means for obtaining 208 a context of the secondary user 312 associated with the computing device and means for transmitting 210 context information 304 indicative of the context of the secondary user. In this example, means 210 are configured to transmit the context information 304 directly to the apparatus 100a, but it will be understood that the context information may be transmitted between the computing device and the apparatus via one or more intermediate devices, as appropriate.

[0049] In still other examples not illustrated here, the means 104 of apparatus 100a comprises the functionality of means 208. The apparatus 100a obtains the context information 304 at obtaining means 104 by receiving or determining (in any suitable manner) a context of the secondary user 312 associated with the device 200 and then obtaining context information indicative of the context directly with means 104. In such an arrangement, the device 200 may not comprise means 208. Device 200 also may not comprise means 210, or means 210 may be configured instead to transmit the context or data indicative of the context so as to allow the means 104 of the apparatus 100a to obtain the context information 304.

[0050] The context information is used to provide context aware adaptation of the volume of the audio 300. Similar to the discussion above with respect to apparatus 100a, the context can be obtained directly by the computing device 200, i.e., using one or more sensors on the computing device and / or can be obtained from one or more other devices, such as associated wearable devices. Additionally or alternatively, the computing device can use information about executing applications on the device or any other suitable context signals to determine a context of the user. In other examples, the context may include reactions of the secondary user that indicate (dis)satisfaction with the audio leakage, such as actions representing irritation or attempts to mitigate the sound (this may include e.g., sending a message to the primary user who is actively listening to the audio, putting on noise cancelling headphones, etc.). The context can be obtained directly by determining the context at the computing device 200, or it can be obtained from another device (for example, user context may be determined / stored in a cloud processing device or a central home hub device). The context information can be transmitted directly to the audio outputting apparatus or to another device, such as a cloud processing device or another device which connects directly with the audio outputting apparatus. The computing device 200 can be a computing device associated with the user. For example, the computing device can be a smartphone or tablet, or other mobile computing device. In other examples, the computing device can be a smart speaker. In some examples, the computing device can be part of a group or smart speakers (e.g. a smart speaker system / network) or collection of user devices, which cooperate to detect the audio and determine / obtain the context. For example, separate devices might provide the volume and context information to the audio playing apparatus.

[0051] In some specific examples, the context comprises a current activity of a user and / or an expected future activity of a user. The context can indicate whether the secondary user 312 is performing a different activity instead of listening to the audio 300 and optionally what that activity is. Some example activities include: reading, working / studying, sleeping, or listening to other audio, playing music, playing a video, or attending a meeting. The current activity can be an activity that the user is performing at a current point in time (e.g. listening to music, watching a film, sleeping, attending a meeting, exercising) and can be determined based on information or context signals from the computing device 200 itself and any applications which are executing on the computing device, as well as information from other devices associated with the secondary user such as e.g. wearables. The expected future activity of the user can be an activity that the user is expected to perform at a later point in time and can be determined based on e.g. calendar events, alarm settings, etc. For example, if a user has an alarm set for early the next morning, or has an early meeting or calendar event, it might be predicted / expected that the user is going to wake up early; therefore, if it is late at night, it might be determined to be that the user is trying to sleep and therefore the context can be indication that the user is likely "sleeping". The current / future activity can include, but is not limited to: sleeping, playing music, playing a video, or attending a meeting.

[0052] To determine the context of the user 312 associated with the computing device 200, a "context detection" rule / algorithm / process can be used. The context detection algorithm requires contextual data. In the simplest case, the proposed approach leverages the set of sensors available on the computing device (which can be e.g., a mobile device, wearable, smart speaker). These sensors can be an inertial measurement unit (IMU), global positioning system (GPS), optical sensors (such as photoplethysmography, PPG, sensors), and / or microphones or other audio sensors. For more advanced context detection, other connected data sources such as digital calendars, scheduled alarms and stored historical activity patterns can additionally be used to help determine context (and potentially higher levels of user stress or susceptibility to stress - e.g., caused by a current or future meeting / deadline / etc.). The computing device may also be connected to other personal devices (e.g., the computing device may be a smartphone connected with a smartwatch and / or earbuds and / or other wearables; the computing device may be a smart speaker connected to a network of smart speakers), which may provide additional contextual information to the computing device and / or apparatus 300. The context can also comprise, for example, whether or not a user is actually in a room in which the audio is detected. For example, presence detectors may be used to detect a presence (or determine an absence) of a user. In other examples, proximity between different devices associated with the secondary user can be used to determine of the secondary user is in the room in which device 200 detects the audio. Periodically, or whenever there is a change in the volume of the speaker, this (optional discovery and) context recognition / determination process can repeat in order to adapt to a potentially ever-changing context or audio volume. Newly detected individuals' devices can go through the full context recognition process described above.

[0053] Individuals' devices that were previously detected do not need to complete the full context determination process; instead, the prior known information about the context can be used in order to preserve the battery of the devices. The activated sensors will only need to determine whether there has been a substantial change in the context from the previously detected one, rather than re-determine the context entirely. For example, this change could be whether the computing device 200 is in the same location / position as before, or whether a sensor / IMU has started counting steps, which might indicate that a user has finished a desk-based activity. Ultimately, this approach allows to minimize power consumption used for context determination, reducing the need to activate all sensors to calculate the full context from scratch with each iteration of this process.

[0054] In this example of FIG.1A, the apparatus 100a further comprises means for determining 108, based on the received context information, a volume threshold. The volume threshold is a maximum volume level which is considered acceptable for the secondary user's context. For example, a user who is exercising may tolerate more leaked sound than a user who is sleeping. The volume threshold can be determined based on predetermined user preferences, historical data, one or more look up tables, one or more transform functions, etc. Any suitable method of determining the volume threshold can be used. The apparatus 100a further comprises means for generating 110 a signal 306 when the volume of the detected audio exceeds the determined volume threshold. The volume can be compared to the volume threshold, and a signal provided based on the comparison. The generated signal can be indicative of, or include, a "decision" about how to adjust or adapt the audio being output so as to reduce the sound leakage experienced by the secondary user 312. In this example, the generated signal can comprise the volume data and the volume threshold and / or an indication of a difference between the volume and the volume threshold and / or can be a control signal indicative of an adjustment of the one or more audio characteristics, the adjustment determined based on the volume and the volume threshold.

[0055] The apparatus 100a of FIG.1A further comprises means for adjusting 112b the one or more audio characteristics, based on the generated signal, to reduce the volume of audio detected at the computing device below the volume threshold. Reducing the volume can comprise adjusting one or more of: a volume, a bass level, a middle level, a treble level, or an equalization. By adjusting the volume and / or balance of the audio being output, the volume detected at the location of the secondary user 312 can be reduced.

[0056] The means 112b for adjusting the one or more audio characteristics can in some examples comprises means for adjusting the audio characteristics based on historical volume data associated with the computing device. For example, there may be one or more functions or transforms which relate a volume at the master audio outputting device and a volume at a location of the device, and the adjustment may be made using these. The adjustment made by the apparatus 100a, and resulting detected volume, may be fed back and used to update these functions or transforms to account for changes in sound leakage within the environment over time (as speaker quality deteriorates, furniture is moved, etc.). In other examples, the feedback may simply be used as part of an iterative process, where the audio characteristics are iteratively adjusted until a desired volume threshold is reached at the computing device.

[0057] In some other examples, the apparatus 100a of FIG.1A comprises means for providing 114 a user notification, at the apparatus, based on the generated signal, the notification indicating that the volume of audio detected at the computing device exceeds the volume threshold. The notification could additionally or alternatively be provided at another computing device associated with the primary user. The notification may be an audible and / or visual alert to the primary user 310. The notification can be instead of or as well as an automatic adjustment by the means 112b. The notification can be provided at the apparatus (whether audibly or visually) and / or at another computing device (such as e.g., a smart phone or tablet). The notification may for example prompt the user change the audio characteristics to reduce the detectable volume, or may notify them that they can change the audio characteristics to increase the detectable volume because it remains below the volume threshold. In some examples, the notification can be provided / output / displayed for a threshold period of time, after which the apparatus may automatically reduce the volume based on the generated output.

[0058] The proposed approach allows for the use of e.g., mobile devices, audio sensors, or computing devices to (optionally automatically) adjust the audio characteristics of audio being output from e.g. a speaker of an audio outputting apparatus (whether it is a television or TV, a radio, a smartphone, a smart speaker, a smart TV, etc.) based on the context of all people (secondary users) who are indirectly exposed to the audio, as opposed to just the primary listener(s). The approach distinguishes between two main actors: the primary listeners / users who are actively consuming the audio content (i.e., the ones in control of the audio playback) and the secondary users, who are indirectly exposed to the audio but are not actively listening to it. The approach then uses the context of those secondary users to dynamically decide how to adapt or adjust the audio characteristics of the output audio to reduce sound leakage. In this way, audio / sound leakage may be dynamically reduced or minimized.

[0059] With reference to FIG. IB, a second example of context-aware adaptive control of sound leakage will be described. The second example includes a system 400 comprising an apparatus 100b, an audio outputting apparatus 150 and a secondary device 200 which is remote from the apparatus 100b. Although shown here as a single device, device 200 can be implemented as one or more devices, as required. The audio outputting apparatus 150 is also called herein a "master" device, since it is the device at which the audio playback can be controlled.

[0060] Operation of the apparatus 100b and computing device 200 is substantially as described above with reference to FIG.1A. In particular, apparatus 100b can perform the operations discussed above with regards to means 102, 104, 108 and 110. However, since the apparatus 100b is not the master device, the means 112a, 114, 112b are shifted from the apparatus to the master device 150, as shown in FIG. IB. Instead, apparatus 100b comprises means for providing 106, to another device configured to output audio with one or more audio characteristics, the generated signal. In this specific example, apparatus 100b comprises means for providing 106 the generated signal 306 to the master device 150 (which is configured to output the audio 300 with one or more audio characteristics).

[0061] The generated signal can comprise a control signal indicative of an adjustment of the one or more audio characteristics, the adjustment determined based on the volume and the volume threshold. The control signal can cause the master device to adjust the one or more audio characteristics, so as to reduce the volume of the detected audio. This generated control signal can comprise one or more commands which cause the master device 150 to adjust the one or more audio characteristics based on the generated signal to reduce the volume of audio detected at the computing device below the volume threshold. This adjustment at the master device is as described above with respect to adjusting means 112b. The generated control signal 306 can additionally or alternatively comprise one or more commands which cause the master device 150 to provide a user notification based on the generated signal, the notification indicating that the volume of audio detected at the computing device exceeds the volume threshold. This provision of a notification is as described above with respect to providing means 114.

[0062] In other examples, the generated signal can comprise the volume and the volume threshold. In this case, some additional processing may be performed by the apparatus 100b and / or master device 150 to compare the volume and volume threshold and cause adjustment of the audio characteristics in dependence on the comparison. Additionally or alternatively, the generated signal can comprise an indication of a difference between the volume and the volume threshold. In this case, some additional processing may be performed by the apparatus 100b and / or master device 150 to cause adjustment of the audio characteristics in dependence on the difference.

[0063] In the example of FIG. IB, apparatus 100b can be considered as an intermediary device, which can sit between the device 200 and the master device 150. The apparatus can be a cloud based data processing device or a home hub type device which does not actively output audio data but can cause another device (e.g. an audio outputting device 150) to adjust the audio characteristic(s) based on the generated signal. In some other examples, the apparatus can be a device which connects directly with the other, audio playing, apparatus (such as a casting device or set top box, for example). In some examples, the apparatus is a smart speaker. The apparatus 100b can thus be implemented in examples as a cloud processing device, a home hub, or part of a smart speaker network which is configured to interface with the device 200 and master device 150. The device and master device can also be implemented as smart speakers; in this arrangement, system 400 can correspond to a network of speakers within a building / house or environment. Apparatus 100b can be any suitable type of device which is able to perform the operations described herein, including a plug-in device that is connected to the master device 150 (such as e.g., a casting device or set top box).

[0064] With reference to FIG.1C, a third example of context-aware adaptive control of sound leakage will be described. The third example includes a system 400 comprising an apparatus 100b, an audio outputting apparatus 150 and two computing devices 200a, 200b. The audio outputting apparatus 150 is also called herein a "master" device, since it is the device at which the audio playback can be controlled.

[0065] Operation of the apparatus 100b and master device 150 is substantially as described above with reference to FIG. IB. However, in this example, the operations performed by device 200 in FIG. IB are distributed between more than one computing device. Optionally, device 200 of system 400 can be implemented as one or more devices, and in the specific, non-limiting, example of FIG.1C the device 200 is implemented as two devices 200a, 200b.

[0066] In this specific arrangement, the volume detection and context information determination are distributed across two different devices. Device 200a implements means 208, 210 described above with reference to FIG.1A, and device 200b implements means 202, 204, 206 described above with reference to FIG.1A. In other examples not illustrated here, the means 204, 208 may be implemented on apparatus 100b, as described above, and means 206, 210 may be removed other otherwise configured to allow the means 102, 104 to obtain the volume data and context information.

[0067] Optional means 212 can be implemented on either device, or on another device which is part of the one or more devices (not shown here). It will be understood that any suitable devices may be used for devices 200a, 200b (and may be used for any other devices not shown here). Devices 200a, 200b may be different types of device, or a same type of device. For example, they may each be a smart speaker within a network 400 of smart speakers. In another example, the context information may be obtained by a mobile or personal computing device 200a (such as a wearable device, a smartwatch or smart phone) and / or the volume detection can be performed by a mobile or static device (such as a static smart speaker, a smart phone, a tablet or person computer, a microphone or other audio sensor, etc.).

[0068] With reference to FIG. 1A, IB and 1C, it will be understood that the means 202, 204, 206, 208, 210 may be distributed across one or more devices 200, 200a, 200b in any suitable manner and / or may be removed entirely. Any suitable number of devices may be used, or any suitable type.

[0069] With reference to FIG. ID, a fourth example of context-aware adaptive control of sound leakage will be described. The fourth example includes a system 400 comprising an audio outputting apparatus 150 and a device 200. Though not shown, the device 200 can be implemented as one or more computing devices, such as discussed above with reference to FIG.1C. The audio outputting apparatus 150 is also called herein a "master" device, since it is the device at which the audio playback can be controlled.

[0070] Operation of the master device 150 is substantially as described above with reference to FIG. IB and 1C. In particular, the means 112a, 114, 112b are implemented at the master device 150, as shown in FIG. ID. However, instead of the apparatus 100a or 100b performing the operations to determine a volume threshold and generate an output, these operations are performed by the device 200 itself. Device 200 is here implemented as a computing device of any suitable type, comprising at least one processor and memory. The device can comprise means 202, 204, 208, 212 as discussed above with reference to FIG.1A, but the device 200 also has additional functionality that was previously described with reference to the means 108, 110 of the apparatus 100a, 100b. In the example of FIG. ID, the computing device 200 further comprises means for determining 214, based on the obtained context information, a volume threshold. Operation of means 214 can be the same as the operation of means 108 of apparatus 100a, described with reference to FIG.1A. In this specific example, the computing device 200 associated with secondary user 312 can measure the amount of leaked sound and combine this information with the obtained context to determine whether the detected volume is acceptable or not. Suitable devices for this task can be any form of computing device that is equipped with a microphone or other means for detecting 202 the audio and one or more processors. The computing device can be, for instance, a smartphone, smartwatch, earbuds, smart glasses, smart speaker, laptop or personal computer, etc.

[0071] The computing device further comprises means for generating 216 a signal 306 when the volume of the detected audio exceeds the determined volume threshold.

[0072] Operation of means 216 can be the same as the operation of means 110 of apparatus 100a, described above with reference to FIG.1A. The generated signal 306 can comprise the volume and the volume threshold. Additionally or alternatively, the generated signal 306 can comprise an indication of a difference between the volume and the volume threshold. Additionally or alternatively, the generated signal 306 can comprise a control signal indicative of an adjustment of the one or more audio characteristics, the adjustment determined based on the volume and the volume threshold. The generated signal 306 includes and / or is based on the volume threshold.

[0073] Computing device 200 optionally further comprises means for transmitting 218 the generated signal. Means 218 can be configured to transmit the generated signal to the master device 150, either directly or via one or more intermediary devices (not shown). The master, or audio outputting, device 150 in this example comprises means for receiving the generated signal 306. This generated signal can cause the master device 150 to adjust the one or more audio characteristics based on the generated signal to reduce the volume of audio detected at the computing device below the volume threshold. This adjustment at the master device can be as described above with respect to adjusting means 112b. The generated signal 306 can additionally or alternatively cause the master device 150 to provide a user notification based on the generated signal, the notification indicating that the volume of audio detected at the computing device exceeds the volume threshold. This provision of a notification can be as described above with respect to providing means 114.

[0074] EXAMPLE 1 : Jane is watching a loud movie on the TV, and Bob is trying to study for an exam the next day. The high volume of the TV is creating a high level of sound leakage, thus disturbing Bob. Bob's computing device (comprising means 202, 204, 208, 214, 216) will, therefore, deem the volume to be unacceptable based on the context. This decision is then output and relayed to the TV (master device) via signal 306. The TV then performs some mitigation actions based on the output, such as adjusting the emitted audio volume or adjusting other audio characteristics (via means 112b, 114) to regulate the volume detected by Bob's computing device. EXAMPLE 2: Jane is watching a loud movie on the TV, and Bob is trying to study for an exam the next day. The high volume of the TV is creating a high level of sound leakage, thus disturbing Bob. An audio sensor associated with Bob (means 202) detects the audio and sends an indication of the detected audio to an apparatus (the TV or another intermediate device). This apparatus (comprising means 102, 104, 108, 110) will, therefore, deem the volume to be unacceptable based on the context. The TV then performs some mitigation actions based on the decision that the volume is unacceptable, such as adjusting the emitted audio volume or adjusting other audio characteristics (via means 112b, 114) to regulate the volume detected by Bob's computing device.

[0075] With reference to FIG.2, an example flow chart illustrates method 2000. The operations of method 2000 can be performed by any suitable apparatus / device or combination of apparatus and / or devices, as discussed above with reference to FIG. 1A-1D.

[0076] Operation S20 optionally comprises detecting audio, wherein the audio is output (operation S10, see FIG.3) with one or more audio characteristics. As discussed above, the audio is output to a primary user 310 (i.e. user who is listening to the audio). The audio is output by an audio outputting device or master device, such as apparatus 100a, 150 discussed above. The audio is detected by a device / apparatus which is remote from the audio playing apparatus. Operation S40 comprises obtaining volume data indicative of a volume of the detected audio. Operation S60 comprises obtaining context information indicative of a context of a user. The user can be a secondary user 312, i.e. a user who is not actively listening to the audio. S80 comprises determining, based on the context information, a volume threshold. S100 comprises generating a signal when the volume of the detected audio exceeds the determined volume threshold. For example, the volume of the detected audio can be compared to the volume threshold. Operations S40, S60, S80, S100 can be performed by any suitable device or combination of devices.

[0077] Method 2000 optionally comprises an additional operation S120 of adjusting the one or more audio characteristics based on the generated signal to reduce the volume below the volume threshold. Operation S120 can be performed by the audio outputting apparatus or master device. Additionally or alternatively, method 2000 can comprising providing a notification indicating that the volume of audio detected at the computing device exceeds the volume threshold. The notification is a user notification, and can be provided to the primary user, either at the audio outputting apparatus or at another device. As shown in FIG.2, the method 2000 can be part of iterative process. After the audio characteristics are adjusted at operation S120, the volume may again be detected and compare to a volume threshold. The characteristics may therefore be iteratively adjusted until the volume threshold is reached. Additionally or alternatively, the detected volume may be used as part of a feedback process to e.g. adjust or update the generated output / para meters associated with the adjustment of the audio characteristics. The process can be repeated periodically, or in response to a change in a volume of the audio output and / or a change in user context.

[0078] Aspects of method 2000 can be understood further with reference to FIG.3. The hardware arrangement of FIG.3 is based on a further example implementation, in which an intermediary apparatus 100c interfaces between the device 200 and the apparatus 100a, 150 for some communications. Apparatus 100c can implement a sound related context detector and performs context related operations (and optionally device detection or "discovery" operations). Device 200detects the audio as in operation S20 (not shown here).

[0079] Apparatus 100a, 150 can operate as the master device (also termed the controller device) and can begin to output audio (S10). The audio can be output locally, i.e. by the audio outputting apparatus 100a, 150 itself (via means 112a). In other examples, the means for adjusting 112b the audio characteristics can be implemented on apparatus 100a, 150, and the apparatus can cause audio to be output from a device remote from the apparatus (i.e. can cause a remote audio outputting device with means 112a to output the audio). Context of the user can be obtained (S60) by apparatus 110c by way of the "context detection" process discussed herein, and in response a signal can be generated (S100) indicating that the volume of the audio detected at device 200 exceeds the volume threshold determined based on the context. This output can be the "feedback about sound leakage level and acceptability" of FIG.3 (i.e. device 200 and / or apparatus 100c perform operations S40, S80). In other examples, this "feedback" can comprise the volume data or information about the volume of audio detected at the device 200, and the obtaining of the volume data and context information (S40, S60) and subsequent determination of the threshold and generation of the output (S80, S100) can instead be performed by apparatus 100a, 150. Either way, the apparatus 100a, 150 then adjusts (S120) the audio output based on the feedback.

[0080] Example implementations In a domestic setting, smart speakers and smart TVs are becoming more and more popular. Many of these are paired / connected to a smartphone or other personal computing device associated with an individual user and / or to a control device of a wider smart home system (which control device may be shared amongst multiple users).

[0081] The above-described context-aware adaptation of volume can be implemented as a software solution that can be deployed in some specific (and non-limiting) examples to 1) devices in a smart home system or 2) another set of connected devices (e.g. where devices that participate in the above-described volume regulation process have the software installed. This can be enforced within a household, or could happen transparently if the software were to be provided as part of the OS of the respective devices).

[0082] These example implementations are now discussed.

[0083] Approach 1 : Implementation on devices in a Smart Home System

[0084] A smart home system functions to manage or control devices in a home, including smartphones or other personal computing devices, TVs, speakers, cameras, lightbulbs, thermostats, and even Wi-Fi routers, among others. These devices which are managed or controlled by the smart home system are called herein "smart devices". In these systems, typically the users can manually connect their devices to the smart home system a single time. The management of the connected smart devices in the smart home system can then be performed by a "control device" for the smart home system; the control device can be a single device, distributed across multiple devices, or implemented at a remote server.

[0085] Usually, each smart device is registered with the smart home system under a personal account, meaning the owner of each device is known to the smart home system. This setup is advantageous for managing volume levels for all residents of a household, and allows for the obtaining of context for each respective user. Following this approach, only the control device for the smart home system needs the above- mentioned software to be downloaded / deployed in order to implement the context- aware volume adaptation. Optionally, the control device discovers all connected smart devices in the smart home system via the home Wi-Fi network or other connections. The control device typically has broadcast functionality, meaning that it can send a message to all connected smart devices to obtain e.g. context information for the users associated with the devices and / or volume data indicative of a volume of audio detected at the device, as appropriate for the particular smart device. These connected devices are the "nearby" devices discussed above. In other examples, there is no such central "discovery" and instead one or more connected devices can initiate the volume regulation process.

[0086] The smart home system implementation is a centralized approach where participating nearby devices send context-related data to the control device, which in turn runs a context recognition algorithm, determines the volume threshold based on the context, and then causes one or more audio characteristics of the output audio (such as volume) to be adjusted based on a comparison of the obtained volume data and determined volume threshold, so as to reduce or minimize sound leakage by regulating the volume detected at the location of a secondary user.

[0087] For example, the control device for the smart home system comprises: means for obtaining 102 volume data 302 indicative of a volume of audio 300 detected at a device (optionally detected at the control device itself); means for obtaining 104 context information 304 indicative of a context of a user associated with the device, which context information can be obtained from one of the smart devices of the smart home system; means for determining 108, based on the context information, a volume threshold; and means for generating 110 a signal 306 when the volume of the detected audio exceeds the determined volume threshold. Regulating the detected volume can comprise providing 106 the generated signal to a remote audio outputting device when the volume of the detected audio exceeds the determined volume threshold, thereby to cause the audio outputting device to adjust the one or more audio characteristics or (where the control device is outputting the audio itself) directly adjusting 112b the one or more audio characteristics to reduce the volume of audio detected at the device below the volume threshold.

[0088] With regards to power consumption, smart home system control devices often do not use a battery and instead must be connected to a power source at all times in order to work. Energy consumption and / or battery life is, therefore, not necessarily a concern when running the context detection algorithm on the control device of a smart home system. Moreover, since the control device of a smart home system can be implemented as a single device within a single home, user data may not need to be transmitted outside of the home to operate the process described herein.

[0089] In one specific implementation, contextual data is sent from one or more connected devices (examples of device 200 in FIG. 1A-1C) to the smart home system control device (example of apparatus 100a, 100b in FIG. 1A-1C). The smart home system control device can access any of connected smart devices (lightbulb, Wi-Fi router, smartphone etc.) to facilitate context detection. The contextual data is transmitted from the connected devices to the control device of the smart home system in any suitable manner, using any suitable wired or wireless connection type or protocol. The same connected devices which provide the context to the control device may also provide the volume data, or the volume data and context can be detected or provided by different devices. Approach 2: Implementation on other connected devices

[0090] In this second example implementation, software is either provided as a standalone application or integrated as part of the Operating System (OS) running on each of the devices participating in the volume regulation process described above (so called "participating devices"). These participating devices can be interconnected to one another (e.g. with a mesh type topology) or each connected to a central participating device (e.g. with a spoke and hub type topology). The participating devices can be connected in any suitable manner.

[0091] An audio playing apparatus 100a or master device 150 outputs audio to an environment. In one specific implementation, the apparatus 100a or master device 150 (which are one of the "participating devices") optionally discovers other nearby "participating devices". The discovery of nearby devices may take place by means of any wired or wireless technology (e.g., Bluetooth, BLE, Wi-Fi, Zigbee, etc.) and / or any other proximity detection technique and enables the master device to establish a connection with all the nearby participating devices. In other examples, the "discovery" is performed by another one of the participating devices. In still other examples, there is no such central "discovery" and instead one or more participating devices can initiate the volume regulation process. Following the discovery phase or the initiation of the volume regulation process, each participating device (example of device 200 in FIG. ID) detects the output audio, runs the context detection algorithm and generates a signal if the volume of the detected audio exceeds a threshold determined based on the context. This signal can then sent back to the audio playing apparatus (i.e. by means 218 for transmitting the generated signal).

[0092] For example, each participating device comprises: means for detecting 202 audio 300, the audio output by the master device 150 or audio outputting apparatus 110a with one or more audio characteristics; means for obtaining 204 a volume of the detected audio; means for obtaining 208 context information indicative of a context of a user associated with the device; means for determining 214, based on the context information, a volume threshold; and means for generating 216 a signal 306 when the volume of the detected audio exceeds the determined volume threshold. By running the context algorithm entirely locally on the participating devices, it is possible to ensure that all the sensitive data and information related to the users' context is kept private. Optionally each participating device further comprises means for transmitting 218 the generated output. The generated signal can be transmitted directly to the remote audio outputting apparatus 100a, 150 or to another, intermediate, device (such as a cloud processing device, home hub, or device connected to the audio playing apparatus. The audio outputting apparatus can be configured to receive the generated output (which can be the volume threshold information or other information which is based on the volume threshold information) and adjust the audio it is outputting accordingly.

[0093] One implementation approach is to integrate the software into the OS, such that a user does not have to e.g. manually download the software as an application onto each participating device. Users without an OS that contains the software might still manually download a software application which would be compatible with the OS implementation. By providing a network of connected participating devices that can implement the context detection and volume determination processes described herein, the processing steps can be distributed across multiple participating devices, allowing for balancing of power consumption and available battery life across the different participating devices.

[0094] As an alternative implementation, only a single participating device ("first participating device") needs the software (as opposed to all participating devices having it either as an app or as part of their OS). Such an implementation would be similar to Approach 1 above, where the first participating device can be any suitable device or apparatus (which could be a smartphone, smart speaker, or other audio device). In such an example, audio detection can be performed at one or more of the other participating devices, which are remote from the first participating device. The first participating device is thus an example of apparatus 100a, 100b in FIG. 1A-1C above. For example, the first participating device comprises: means for obtaining 102 volume data 302 indicative of a volume of audio 300 detected at one or more of the participating devices; means for obtaining 104 context information 304 indicative of a context of a user associated with the participating device at which the audio is detected; means for determining 108, based on the context information, a volume threshold; and means for generating 110 a signal 306 when the volume of the detected audio exceeds the determined volume threshold. The volume data and context information can be obtained via one or more of the other participating devices. In some examples, the first device is the audio playing apparatus 100a, 150 and comprises means for outputting 112a the audio with one or more audio characteristics. In other examples, another one of the participating devices can be configured as the audio playing apparatus and the first device acts as an intermediary device, such as apparatus 100b. Such a centralized solution can reduce power consumption and / or battery drain at the other participating devices, which can be beneficial in implementations where the participating devices are mobile computing devices (such as e.g. smartphones).

[0095] Discovery of nearby devices

[0096] The optional step of "discovery" of nearby devices is now discussed in more detail.

[0097] The discovery can be performed by any suitable device or apparatus (such as apparatus 100a, 100b, 100c discussed above). In some implementations, such as when the volume adaptation process is initiated by the device 200 associated with the user, there may be no such "discovery step".

[0098] The discovery step can be performed once for each system of connected devices, and a list or details about the different connected devices can be stored for use in implementing the volume regulation process described herein. In some examples, discovery can be performed periodically so that an updated list of connected devices may be performed. In some other examples, discovery may be performed each time the audio outputting apparatus is operated to play audio. In some implementations, discovery that a device (such as device 200) is near to the audio outputting apparatus may cause said device to detect audio which is output by said audio playing apparatus; for example, the discovery process may trigger the device to start detecting audio. With reference to the specific implementations discussed above, in Approach 1 (the smart home system approach), discovery of nearby devices can be easily done because each device automatically connects to the smart home system via any suitable wired or wireless connection. The control devices of smart home systems already have the technology implemented to recognize and contact the connected smart devices within the smart home system. Any suitable connection techniques or protocol can be used.

[0099] For Approach 2 (the connected devices approach), the discovery of nearby participating devices can be done using various conventional techniques. For example, devices being connected to the same Wi-Fi network can be used to infer that they are located nearby to one another. Similarly, Bluetooth-enabled devices can discover each other; given the limited range of Bluetooth signal, all discovered devices can be considered nearby devices. NFC is a short-range wireless communication technology which is commonly used e.g. for contactless payments, but it can be used to detect nearby devices having NFC tags or readers. Ultrasonic waves can be sent and received by devices (including smartphones and smart speakers) at inaudible frequencies, and can be used to determine that devices are nearby to one another, and even the proximity of such devices. Global positioning systems (GPS) can also be used to determine location of devices with high accuracy. Indoor positioning systems can be used instead of or as well as GPS e.g. in an indoor or urban environment where GPS lacks precision. Cellular network connectivity can also be used. Cellular network infrastructure can track location using the signal strength and tower connectivity of connected devices.

[0100] Audio detection

[0101] The primary listener is the person or persons who plays the audio at the audio outputting apparatus and wants to "consume" it (i.e., is actively listening to it). A user who can hear the audio as background noise, but is not actively consuming it, is a secondary user 312.

[0102] Notably, the background or ambient noise in the environment of the secondary user can be affected by factors other than the audio being output by the audio outputting apparatus. The ambient noise that is detected by the detecting means 202 of device 200 therefore might not correspond to the audio output from the audio outputting apparatus, and might therefore not be representative of sound leakage. Determining the volume of the output audio detected at the device 200 can involve the extraction of a component of the detected audio that corresponds to the audio output and determining the volume of only this component.

[0103] This can be done in various ways. In one example, the audio output can be compared with the sound detected at the device 200 to determine what component of the detected audio is correlated with the audio output from the audio playing apparatus.

[0104] In another example, information about e.g. certain frequencies of the audio output can be determined and the volume can be determined only at a particular frequency or set of frequencies that correspond to content in the audio output. In another example, volume measurements taken at particular times could be compared; if e.g., the audio outputting device is quiet at time X and loud at time Y, then an overall difference in volume between those two times can be determined. In other examples, the device can receive an indication of the audio being output by the audio outputting device, wherein determining a volume of the detected audio comprises determining a volume of the detected audio based on the received indication. This indication can be received from the audio outputting apparatus itself (e.g. from apparatus 100a or 150), or from an intermediary apparatus (such as apparatus 100b, 100c).

[0105] Other approaches can use historical knowledge of the relationship between a certain set of audio characteristics from the audio outputting device and the audio detected at the location of the device 200. Based on past instances, if the volume detected by the device 200 does not match historical relationships between the speaker volume and the detected noise level, then there could be a third party noise source. The historical relationship can be used to isolate the audio output from this third party noise source.

[0106] Context detection

[0107] The context of the secondary user is used to determine whether the audio characteristics of the audio being output by the audio outputting apparatus should be adjusted. The context can be determined in any suitable manner, and based on any suitable signals.

[0108] The context detection can be run at any suitable device or apparatus of system 400, as discussed above. Detection of audio by the means 202 (such as at device 200), i.e., operation S10 of FIG.2, may cause context detection to be initiated. In other examples, the context detection is performed independently of the audio detection.

[0109] The context information can optionally be stored and can be provided for use in the volume regulation process, as required. In one specific implementation, the context is obtained by running a context detection algorithm. The context detection algorithm takes as input a variety of contextual information associated with the secondary user to determine the context of the secondary user. The context detection algorithm can use information from sensors of a mobile computing device of the secondary user to get contextual information. These sensors include but are not limited to IMUs, GPS, Wi-Fi sensing localization, Bluetooth, PPG, and microphones. The context detection algorithm can additionally or alternatively obtained information from a personal wearable device, such as a smartwatch, which is likely to have more accurate and fine-grained sensor readings relating to a biological or physiological condition of the user (e.g., to determine stress level, activity, etc.).

[0110] Additional information from which a user's likely activity can be inferred can also be obtained from e.g. paired data sources. This additional information can include a calendar appointment relating to a particular activity. This additional information can include an alarm clock for a particular time. In this way, information such as work meeting schedules, or estimating sleeping hours, can be included in the context determination. The algorithm can look to determine relevant context factors such as the following:

[0111] • Social context. The social context can indicate a location of the user and / or whether the user is alone. For example, the social context can indicate: Is the user(s) at home, alone, in a public place, at work, at a restaurant, or is there a baby / toddler napping? • Environmental noise context. The environmental noise context can indicate the presence of ambient noise in an environment of the user. This can include whether there is ambient noise, what type it is, and / or how loud it is. For example, the environmental noise context can indicate: Is the user(s) listening to music (with earbuds or a speaker), or watching TV? Is there ambient noise from road noise, lawnmowers, construction, rain, or cooking? How loud is the noise?

[0112] • Activity context. The activity context can indicate the likely activity being performed by the user at the current time and / or an expected future activity of the user. For example, the activity context can indicate: Is the user(s) working, asleep, driving? Do they have an early meeting scheduled? Do they have any exercise or leisure activities scheduled? All of these factors can form part of the contextual information about the secondary user. The information can be used to determine an overall context of the secondary user. The context algorithm can optionally output a numerical value indicative of the determined context, where the numerical value is associated with an indication of how likely it is that the user will be disturbed by the audio from the audio outputting apparatus. For example, a low value can indicate the user is likely to be disturbed. The context can be used to determine a threshold volume, as discussed above with respect to operation S80. For example, a threshold volume may be lower if the context indicates the secondary user should not be disturbed (optionally corresponding to a low value from the context algorithm), and higher if the context indicates that the audio from the audio outputting apparatus is unlikely to disturb the user.

[0113] Regulation of detected volume

[0114] The next operation relates to how the contextual information is used to mitigate the sound level perceived by the secondary user 312 (sound leakage). There are many possible actions that can alleviate the noise leakage problem.

[0115] One example is to simply reduce the output volume of the audio output or to adjust other audio characteristics (such as a bass level, a middle level, a treble level, a balance or an equalization). Different frequencies will be attenuated at different levels by e.g. walls or other obstacles, so changing the equalization can therefore reduce the disturbance to the secondary user without necessarily reducing the overall volume the primary listener hears.

[0116] Another example is to activate a smart speaker or other device closer to a "disturbed" secondary user to play a noise-cancelling sound (e.g., anti-noise as is used in active noise cancellation), which actions can be initiated based on the signal that is generated when the volume of the detected audio exceeds the threshold volume determined based on the context. These would require different sequences of tasks and communication, such as detecting if there is a candidate device near the secondary user (e.g., with a Bluetooth lookup or other mechanism) and choosing which sound to play based on e.g. the signal and / or one or more frequency components of the audio (e.g., anti-noise or music of an appropriate genre and volume to cover up the leaked sound). In another example, it can be possible to turn on some other non-speaker device which makes noise, such as a fan, air filter, air conditioner, generator, dishwasher, etc. in order to mask audio leakage.

[0117] If the decision is to change the volume / audio characteristics, then this can be communicated to the audio outputting apparatus or master device 150 (which can be e.g. the primary listener's smartphone, a TV, or a speaker of the smart home system) by sending the generated signal 306 via any suitable wired or wireless channel (e.g. Bluetooth, Wi-Fi network, cellular network, ultrasonic audio, etc.). The audio outputting apparatus then adjusts the one or more audio characteristics accordingly, optionally trying to comply with signals 306 transmitted from a plurality of secondary users if applicable, e.g. by complying with the largest adjustment of audio characteristics.

[0118] The software on the audio outputting apparatus can in some examples be implemented as a user-selectable option; the software can thus be turned on / off by the primary listener, such that the primary listener can disable the volume regulation process and have full control over the volume. In some examples, the volume regulation process may be enabled but automatic adjustment of the audio characteristics may be turned off; the signals 306 might in this instance cause outputting of a notification to notify the primary listener that they are actively disturbing people in their surroundings.

[0119] In some examples, feedback may be gathered to improve the adjustment of the audio characteristics over time. For example, the secondary user 312 may be asked for input to learn their volume preferences and the contexts in which audio characteristics should be adjusted to reduced leaked audio. Thus, over time two things can be learned: 1) for what contexts should the audio characteristics of the audio output at the audio outputting apparatus be adjusted so as to regulate the volume detected at the device 200, and 2) how much should the audio characteristics be adjusted (e.g. what detected volumes are acceptable to the disturbed user for the context, and / or how does the adjustment correlate to the volume of detected audio at the device 200). With this user input, one or more transforms, algorithms, functions, or models, including machine learning models or simple decision trees, can be used to automatically decide whether the audio characteristics should be adjusted and to learn which detected volumes are disturbing to the secondary user.

[0120] In other examples, historical volume data can be used for adjusting the audio characteristics. For example, the means for adjusting the one or more audio characteristics comprises means for adjusting the audio characteristics based on historical volume data associated with the device. For example, based on a location of the device and / or type of the device, there may be a known transform function relating between the adjustment of the one or more audio characteristics at or by the audio output apparatus and the volume which is detected at the device 200. The transform function can be determined / generated based on historical volume data. In another example, the historical data may be stored in e.g. a database or look up table, and the adjustment can be determined based on accessing the stored historical data. Example Apparatus

[0121] FIG. 4 shows an apparatus according to some example embodiments, which may comprise any of the apparatus 100, 100a, 100b, 100c, 200, 200a, 200b and 150 described herein. The apparatus may be configured to perform the operations described herein, for example operations described with reference to any disclosed process.

[0122] The apparatus comprises at least one processor 4000 and at least one memory 4001 directly or closely connected to the processor. The memory 4001 includes at least one random access memory (RAM) 4001a and at least one read-only memory (ROM) 701b. Computer program code (software) 4005 is stored in the ROM 4001b. The apparatus may be connected to a transmitter (TX) and a receiver (RX). The apparatus may, optionally, be connected with a user interface (UI) for instructing the apparatus and / or for outputting data. The at least one processor 4000, with the at least one memory 4001 and the computer program code 4005 are arranged to cause the apparatus to at least perform at least the method according to any preceding process, for example as disclosed in relation to the flow diagram of FIG. 2 and related features thereof.

[0123] FIG. 5 shows a non-transitory media 500 according to some embodiments. The non- transitory media 800 is a computer readable storage medium. It may be e.g. a CD, a

[0124] DVD, a USB stick, a blue ray disk, etc. The non-transitory media 800 stores computer program code, causing an apparatus to perform the method of any preceding process for example as disclosed in relation to the flow diagrams and related features thereof. A memory may be volatile or non-volatile. It may be e.g. a RAM, a SRAM, a flash memory, a FPGA block ram, a DCD, a CD, a USB stick, and a blue ray disk.

[0125] If not otherwise stated or otherwise made clear from the context, the statement that two entities are different means that they perform different functions. It does not necessarily mean that they are based on different hardware. That is, each of the entities described in the present description may be based on a different hardware, or some or all of the entities may be based on the same hardware. It does not necessarily mean that they are based on different software. That is, each of the entities described in the present description may be based on different software, or some or all of the entities may be based on the same software. Each of the entities described in the present description may be embodied in the cloud.

[0126] Implementations of any of the above described blocks, apparatuses, systems, techniques or methods include, as non-limiting examples, implementations as hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. Some embodiments may be implemented in the cloud. It is to be understood that what is described above is what is presently considered the preferred embodiments. However, it should be noted that the description of the preferred embodiments is given by way of example only and that various modifications may be made without departing from the scope as defined by the appended claims.

Claims

Claims1. Apparatus (100a, 100b, 200), comprising: means for obtaining (102, 204) volume data (302) indicative of a volume of audio (300) detected at a device; means for obtaining (104, 208) context information (304) indicative of a context of a user associated with the device; means for determining (108, 214), based on the context information, a volume threshold; and means for generating (110, 216) a signal (306) when the volume of the detected audio exceeds the determined volume threshold.

2. The apparatus (100a, 100b) of claim 1, wherein the device is remote from the apparatus.

3. The apparatus (100a) of claim 2, wherein the apparatus further comprises means for outputting (112a) the audio with one or more audio characteristics.

4. The apparatus of claim 3, further comprising: means for adjusting (112b) the one or more audio characteristics, based on the generated signal, to adjust the volume of audio detectable at the device.

5. The apparatus of claim 4, wherein the means for adjusting the one or more audio characteristics comprises means for adjusting the audio characteristics based on historical volume data associated with the device.

6. The apparatus of any of claims 2 to 5, further comprising: means for providing (114) a user notification based on the generated signal, the user notification indicating that the volume of audio detected at the device exceeds the volume threshold.

7. The apparatus (100b, 200) of claim 1 or claim 2, further comprising means for providing (106), to another device (150) configured to output audio with one or more audio characteristics, the generated signal, wherein the generated signal comprises one or more commands which cause the another device to: adjust the one or more audio characteristics to reduce the volume of audio detectable at the device; and / orprovide a user notification, the notification indicating that the volume of audio detected at the device exceeds the volume threshold.

8. The apparatus of any of claims 2 to 7, wherein the one or more audio characteristics comprise one or more of: a volume, a bass level, a middle level, a treble level, or an equalization.

9. The apparatus of claim 1, wherein the apparatus (200) is comprised by the device, the apparatus comprising means for detecting (202) the audio (300).

10. The apparatus of claim 9, wherein the audio is output by another device (150) remote from the apparatus, the apparatus further comprising means for receiving (212) an indication (308) of the audio being output by the other device, wherein the means for obtaining the volume data comprises means for obtaining the volume data based on the received indication.

11. The apparatus of claim 9 or claim 10, the apparatus further comprising means for transmitting (218) the generated signal.

12. The apparatus of any preceding claim, wherein the context comprises a current activity of the user and / or an expected future activity of the user.

13. The apparatus of claim 12, wherein the activity comprises: sleeping, consuming media, or attending a meeting.

14. The apparatus of any of claims 1 to 13, wherein the generated signal comprises: the volume data and the volume threshold; and / or an indication of a difference between the detected volume and the volume threshold; and / or a control signal indicative of an adjustment of the audio, the adjustment determined based on the detected volume and the volume threshold.

15. A system, comprising: an apparatus (100a, 150) comprising means for outputting audio (212a) with one or more audio characteristics; a device comprising means for detecting (202) the output audio (300);means for obtaining (102, 204) volume data (302) indicative of a volume of the detected audio (300); means for obtaining (104, 208) context information (304) indicative of a context of a user associated with the device; means for determining (108, 214), based on the received context information, a volume threshold; and means for generating (110, 216) a signal (306) when the volume of the detected audio exceeds the determined volume threshold.

16. The system of claim 15, wherein: the apparatus further comprises means for adjusting (212b) the one or more audio characteristics, based on the generated signal, to reduce the volume of audio detectable at the device; and / or the system further comprises means for providing (214) a user notification based on the generated signal, the notification indicating that the volume of audio detected at the device exceeds the volume threshold.

17. The system of claim 15 or claim 16, wherein the device is a smart speaker and / or wherein the apparatus is a smart speaker.

Citation Information

Patent Citations

  • Equipment volume control method and related device

    CN117014547A

  • Electronic media volume control

    US20170094434A1