Audio notifications
By using machine learning to attune audio notifications to the parameters of ongoing audio content, the integration of notifications is achieved without disrupting the user's experience, enhancing the delivery of key information.
Patent Information
- Application Number
- GB2023017282
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2043-11-10
AI Technical Summary
Existing audio playback devices struggle to seamlessly integrate audio notifications with ongoing audio content without disrupting the user's experience.
An apparatus and method that utilize machine learning models to identify suitable insertion portions in audio content and generate audio notifications attuned to the audio content's parameters, such as tempo, pitch, and language style, allowing the notifications to be seamlessly integrated.
Enables the delivery of key information through audio notifications that blend with the audio content, maintaining the user's experience without interruption.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNOLOGICAL FIELD Examples of the disclosure relate to audio notifications. Some relate to generating audios notifications for combining with other audio content. BACKGROUND If a user is using a device such as mobile phone then the device can provide multiple types of audio outputs. For example, the device could be used to playback audio content such as music or a podcast or other types of audio. The device can also be used to provide audio notifications that alert the user to incoming communications or notifications from one or more applications of the device. BRIEF SUMMARY According to various, but not necessarily all, examples of the disclosure there is provided an apparatus comprising means for: identifying one or more insertion portions of audio content that is for playback to a user; obtaining a communication; extracting information from the obtained communication, wherein the information is extracted using language processing; generating one or more audio notifications based on the extracted information wherein the one or more audio notifications are generated to be in a format that is attuned with the audio content; and adding the one or more generated audio notifications to the identified one or more insertion portions of the audio content to generate a combined output. Identifying one or more insertion portions may comprise determining a voice component within the audio content and identifying one or more portions where the voice component is below a threshold. Identifying one or more insertion portions may comprise determining one or more portions in the audio content where one or more parameters of the audio content are below a threshold. Identifying one or more insertion portions may comprise identifying one or more portions of the audio content that are replaceable with an audio notification. A machine learning model may be used to identify the one or more insertion portions. The means may be for adapting a duration of the one or more insertion portions to correspond to a duration of the generated audio notification. The obtained communication may comprise, an incoming message, a notification from an application, or a calendar notification. The communication may comprise text. A machine learning model may be used to perform the language processing. Attuning the audio notification with the audio content may comprise adapting one or more parameters of the audio notification to correspond to the audio content wherein the one or more parameters comprise at least one of: tempo, pitch, language style, voice style The means may be for enabling the combined output to be played back to a user. The apparatus may be at least one of: a telephone, a smart speaker, a computing device, a gaming device, or a headset. According to various, but not necessarily all, examples of the disclosure there is provided a method comprising: identifying one or more insertion portions of audio content that is for playback to a user; obtaining a communication; extracting information from the obtained communication, wherein the information is extracted using language processing; generating one or more audio notifications based on the extracted information wherein the one or more audio notifications are generated to be in a format that is attuned with the audio content; and adding the one or more generated audio notifications to the identified one or more insertion portions of the audio content to generate a combined output. According to various, but not necessarily all, examples of the disclosure there is provided a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform at least: identifying one or more insertion portions of audio content that is for playback to a user; obtaining a communication; extracting information from the obtained communication, wherein the information is extracted using language processing; generating one or more audio notifications based on the extracted information wherein the one or more audio notifications are generated to be in a format that is attuned with the audio content; and adding the one or more generated audio notifications to the identified one or more insertion portions of the audio content to generate a combined output. While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all of the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all of the features, in any combination, may be implemented by / comprised in / performable by an apparatus, a method, and / or computer program instructions as desired, and as appropriate. BRIEF DESCRIPTION Some examples will now be described with reference to the accompanying drawings in which: FIG. 1 shows an example use case scenario; FIG. 2 shows an example method; FIG. 3 shows an example method; and FIG. 4 shows an example controller. The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Corresponding reference numerals are used in the figures to designate corresponding features. For clarity, all reference numerals are not necessarily displayed in all figures. DETAILED DESCRIPTION Examples of the disclosure relate to apparatus and methods that enable audio notifications to be generated for a user. The audio notifications can be generated so that they blend in with audio content that is currently being played back by the apparatus. The audio notifications can be generated so that they blend in with the current audio content. The audio notifications generated in this way enable an output of key information to be provided without interrupting the audio content. Fig. 1 shows an example use case scenario for examples of the disclosure. In this example a user 106 is using an apparatus 100 to listen to audio content. The apparatus 100 can be any apparatus that enables the playback of audio content and which can receive one or more communications. The apparatus 100 could be a telephone, a smart speaker, a computing device, a gaming device, a headset, or any other suitable type of apparatus. The audio content that is played back by the apparatus 100 can be any type audio content. The audio content can comprise music, speech or any other type of audio. In some examples the audio content can comprise multiple audio items. For instance, the audio content could comprise multiple songs. There could be a gap between respective audio items in the audio content. The apparatus 100 shown in Fig. 1 comprises a controller 102 and an audio playback module 104. Only components referred to in the following description are shown in Fig. 1. The apparatus 100 could comprise other components that are not shown in Fig. 1. For example, the apparatus 100 could comprise one or more transceivers that enable radio communication and / or any other suitable types of components. The controller 102 can comprise any means that can be used to control the functions of the apparatus 100. Examples of a controller 102 are shown in more detail in Fig. 4. The controller 102 can be configured to receive inputs from the audio playback module 104 and / or any other suitable components of the apparatus 100. The controller 102 can be configured to provide outputs to the audio playback module 104 and / or any other suitable components of the apparatus 100. The controller 102 could be configured to implement methods such as the methods of Figs. 2 and 3. The audio playback module 104 can enable audio content to be played back to a user. In some examples the audio playback module 104 can enable audio content to be played back directly. For instance, the apparatus 100 could comprise one or more loudspeakers that can be controlled by the audio playback module 104 to generate sound. In some examples the audio playback module 104 can enable audio content to be played back via one or more peripheral devices 108. The peripheral device 108 could comprise headphones as shown in Fig. 1, one or more loudspeakers or any other suitable device that can generate sound. The audio content can be provided from the apparatus 100 to the peripheral device 108 via a wired or wireless connection. In the example of Fig. 1 a user 106 is listening to the audio content via the peripheral device 108. The apparatus 100 can receive one or more communications and can be configured to provide an audio notification to the user to alert the user 106 to the communication. Fig. 2 shows an example method of generating audio notifications so that they blend in with the current audio content. The method of Fig. 2 could be implemented using an apparatus 100 as shown in Fig. 1 and / or any other suitable type of apparatus. The method comprises, at block 200, identifying one or more insertion portions of audio content that is for playback to a user 106. In some examples the audio content can comprise audio content that is stored in the apparatus 100 or stored in a location that can be accessed by the apparatus 100. This could comprise downloaded or streamed content for example. In some examples the audio content can comprise live content such as a teleconference or a radio or television broadcast or other suitable content received in real time. An insertion portion is a part of the content to which an audio notification could be added. Identifying one or more insertion portions can comprise identifying portions of the audio content which would be suitable for combining with an audio notification or for adding an audio notification to. A suitable insertion portion can be a part of the audio content that can be modified without affecting the user’s 106 experience of the audio content or that causes little disruption to the overall content. The parameters that make a portion of audio content suitable for combining with an audio notification or for adding an audio notification to can be dependent upon the type of audio content. For instance, if the audio content comprises songs then a suitable insertion portion could be an instrumental part or a part which is repeated multiple times or a gap between songs. If the audio content comprises speech then a suitable insertion portion could be a pause in the speech. In some examples identifying one or more insertion portions can comprise determining one or more portions in the audio content where one or more parameters of the audio content are below a threshold. The parameters could be noise or energy levels. In some examples a suitable insertion portion could be a portion of the audio in which there is no speech, or in which there is very little speech. There could be very little speech if there is speech in a background component, or if there are gaps or pauses in the speech. In such examples identifying one or more insertion portions can comprise determining a voice component within the audio content and identifying one or more portions where the voice component is below a threshold. The voice component can be identified using any suitable filtering or processing. The threshold can be used to determine if the speech is a main component or a background component and / or if there are gaps or pauses in the speech. In some examples identifying one or more insertion portions comprises identifying one or more portions of the audio content that are replaceable with a notification. A portion of the audio content can be replaceable if it doesn’t contain information that is key to the understanding or enjoyment of the audio content. For instance, a song could comprise portions that are repeated such as a chorus, and this could be replaced with an alternative chorus. In some examples the audio content can comprise multiple content items. For example, the audio content could comprise multiple songs. A gap or interval could be provided between the songs. An insertion portion could be identified in a gap between the audio items and / or spanning across two audio items. Any suitable means can be used to identify the one or more insertion portions. In some examples machine learning models can be used to identify the one or more insertion portions. At block 202 the method comprises obtaining a communication. The communication can comprise any message or input that comprises information that is to be relayed to the user of the apparatus 100. In some examples the communication can comprise incoming messages or inputs from a source that is external to the apparatus 100. Such communications could be text messages, chat messages, multi-media messages, incoming calls, or any other suitable type of communications. In some examples the communication could comprise an audio or video message that comprises text-based metadata. In some examples the communication can comprise incoming messages or inputs from a source that is internal to the apparatus 100. Such messages can comprise a notification from an application that is run on the apparatus 100. For instance, the communication could be a notification from a calendar application or a health monitoring application or any other suitable type of application. The communication can comprise different components. For example, a text-based message can comprise a content component which comprises the text and some metadata components that indicate the sender of the message, a subject title of the message and any other suitable information. Similarly, a multi-media message could comprise a media component that can comprise an audio or image or video component and some metadata components. At block 204 the method comprises extracting information from the obtained communication. The information is extracted using language processing. The information can be extracted from any suitable part or parts of the communication. In some examples the information can be extracted from the content of the message. For instance, if the message is a text message or email the content of the message or email could be analyzed to extract key information. If the message is a multimedia message such as a video or voice message then speech-to-text processing can be used to generate a text-based component. The information can then be extracted from the text-based component. In some examples the information can be extracted from metadata associated with the message. For instance, the information could be extracted from an indication of the source of the message. This could be used to identify information such as the sender of a text message or the application that has provided a notification. Any suitable means can be used to perform the language processing. In some examples a machine learning model can be used to perform the language processing. At block 206 the method comprises generating one or more audio notifications based on the extracted information. The one or more audio notifications are generated to be in a format that is attuned with the audio content. Attuning the audio notification with the audio content can comprise adapting one or more parameters of the audio notification to correspond to the audio content. The parameters that are adjusted can comprise tempo, pitch, language style, voice style, timbre, and or any other suitable parameter or combination of parameters. The audio notifications can be generated by converting the information extracted at block 204 into an audio output. This can comprise extracting key information from the message. This information can be extracted from any part of the message. The audio notifications that are generated can comprise a spoken or sung output that comprises a summary or indication of the key information from the message or metadata of the message. The words used in the audio notification can be phrased so as to copy the lyrical style of the audio. For instance, if the audio content is a song the wording of the audio notification can be phrased so that the rhythm and timing of the audio notification is the same as the rhythm and timing of the audio content. The audio notifications can be generated to match the sounds of the audio content. For example, the voice used to provide the audio notification can be adjusted so that it sounds similar to, or substantially similar, to one or more voices used in the audio content. The audio notification can be generated to be of a length that fits in to the identified insertion portion. The audio notification can be generated so that it fits in to the identified insertion portion without overlapping onto the audio content. In some examples the duration of the insertion portion can be adapted to correspond to a duration of the generated audio notification. The duration of the insertion portion can be adapted to enable the audio notification to be inserted without overlapping onto the audio content. For instance, the length of a gap or pause in spoken or sung content could be increased or additional bars could be added to an instrumental section of a song. At block 208 the method comprises adding the one or more generated audio notifications to the identified one or more insertion portions of the audio content to generate a combined output. The generated combined output can then be played back to the user 106. Fig. 3 schematically shows an example method that could be used in some implementations of the disclosure. The method could be implemented by an apparatus 100 as shown in Fig. 1 or any other suitable type of apparatus. In the example of Fig. 3 a communication 300 is obtained. The communication 300 can be an external message that is received from another device or apparatus. In such examples the communication 300 could be received using any suitable communications protocols. In some examples the communication 300 could be an internal message that is received from one or more applications implemented by the apparatus 100. In this example the communication 300 is a message that has been received through a messaging or chat application. The message can comprise multiple components. In this example the message comprises a metadata component 302 and a text-based component 304. The metadata component 302 can comprise information that indicates the sender of the message, a subject title of the message, a time of the message and / or any other suitable information. The text-based component 304 can comprise text-based information. The text of the text-based component 304 can be written by the sender of the message or can originate from any other suitable source. The communication 300 is provided to a natural language processing (NLP) module 306. The NLP module 306 can be configured to analyse the communication 300 to extract information from the communication 300. The NLP module 306 can be configured to analyse one or more of the respective components of the communication 300. For instance, in the example of Fig. 3 the NLP module 306 can be configured to analyse the text-based component 304 and / or the metadata component 302. The NLP module 306 can use any suitable algorithms or means to extract the information from the communication 300. The NLP module 306 can use machine learning models and / or any other suitable means. The NLP module 306 can be configured to extract the information that is to be used to generate a personalized audio notification for the user of the apparatus 100. The information that is extracted can relate to the context and content of the communication 300. In some examples the information that is extracted can comprise information such as, a sender or source of the communication 300, a subject to which the communication 300 relates, an identification of any other people identified in the communication 300, an identification of an event or time indicated in the communication, and / or any other suitable information. The NLP module 306 can provide extracted information 308 as an output. The extracted information can comprise a precis of the communication. As an example, the incoming communication 300 could comprise a message with the following parameters: Communication type WhatsApp message Sender Liana Draghescu Time 2:27pm Text-based portion Hi, you’ve been invited to the meeting. It’s today at 5pm. Could you let us know if you can make it? Thanks. The extracted information that is output by the NPL could then comprise the key information from this message. The extracted information could comprise: Liana, meeting, 5pm, can you attend? The extracted information 308 can then be used to generate a personalized audio notification. The audio notification that is generated is specific to the communication 300 that has been obtained because it is based on information extracted from the communication 300. This can cause different communications 300 containing different information to result in different audio notifications being generated. The obtaining of the communication 300 and the extracting of information from the communication 300 can occur while the apparatus 100 is playing back audio content 310. The audio content 310 can comprise any audio that the user 106 of the apparatus 100 could be listening to. The audio content can comprise music, spoken content, and / or any other type of audio content 310 or combinations of audio content 310. In some examples the audio content 310 can comprise multiple content items. There can be gaps or intervals between respective content items. The audio content 310, or at least part of the audio content 310, can be stored in the apparatus 100. In some examples the audio content 310 can be stored temporarily in the apparatus 100 so as to enable playback of the audio content 310. The extracted information 308 and, at least a portion of, the audio content 310 can be provided to an audio notification generator module 312. The portion of the audio content 310 that is provided to the audio notification generator module 312 can comprise a portion that is to be played back within a given time interval, for example it could comprise the audio content that is to be played back within the next minute, or within any other suitable time interval. The audio notification generator module 312 is configured to generate one or more audio notifications based on the extracted information 308 and combine this with the audio content 310. The audio notification generator module 312 can comprise multiple sub-modules. In the example of Fig. 3 the audio notification generator module 312 comprises three sub-modules 314, 316, 318. In this example the first sub-module is an insertion portion sub-module 314, the second sub-module is an audio notification creation sub-module 316 and the third submodule is a combination sub-module 318. Other types or combinations of sub-modules could be used in other examples. The insertion portion sub-module 314 is arranged to analyze the audio content 310 and identify one or more insertion portions. The insertion portions can be identified so as to enable one or more audio notifications to be combined with or added to the audio content 310. The identified insertion portions can be a part or parts of the audio content 310 that are suitable for being combined with the generated audio notification. In some examples the insertion portion can comprise a part or parts of the audio content 310 that can be modified without affecting, or without substantially affecting, the information that is conveyed by the audio content 310. In some examples the insertion portion can be a portion of the audio content 310 that can be replaced with an audio notification. For example, they can comprise portions that don’t comprise any speech or words or otherwise convey significant information. In some examples they can comprise portions that are repeated. In some examples the insertion portion could be a part of the audio content 310 that has reduced voice components compared to other parts of the audio content 310. For instance, if the audio content 310 is a spoken audio content an insertion portion could be a pause or interval between speakers. If the audio content 310 is music an insertion portion could be an instrumental part. In some examples an insertion portion could comprise a part of the audio content 310 that comprises voice components. The parts with voice components that are suitable for use as an insertion portion could be components that are repeated within the audio content, for example, the chorus of a song. In some examples the audio content 310 can be adapted to create an insertion change the duration of an insertion portion. For instance, if the audio content 310 comprises people speaking, a pause or gap could be added at an appropriate point or an existing gap could be extended. If the audio content 310 comprises music or people singing then an instrumental portion could be added or extended so as to create a suitable insertion portion. The audio notification creation sub-module 316 is arranged to convert the extracted information 308 into an audio notification. The audio notification creation sub-module 316 uses the audio content 310 to convert the extracted information 308 into an audio notification. The audio notification creation sub-module 316 can determine stylistic parameters that can be replicated or imitated in the audio notification. The stylistic parameters could be the meter of lyrics or spoken components. For instance, the meter or number of beats in a line of lyrics could be replicated in the audio notification so that the audio notification is attuned to the audio content 310. In some examples the stylistic parameters could be the type of language that is used. Different types of language could use different vocabularies and different sentence structures. The type of language used could be formal language, informal language, regional dialects, or any other suitable type. The language type that is used to select the vocabulary and sentence structure for the audio notification can be based on the language type that is used in the audio content 310. In some examples, the stylistic parameters could be audio parameters that affect how the audio notification sounds. For example, a voice used to provide the audio notification can be adjusted so that it sounds similar to, or substantially similar, to one or more voices used in the audio content 310. Other stylistic parameters or combinations of stylistic parameters could be used in other examples. In some examples user preferences can be taken into account to select the stylistic parameters and / or the insertion portions that are selected. For instance, in some examples a user 106 could indicate a preference for a particular style of the audio notifications. They could indicate that they prefer formal or informal language or prefer female or male voice or select particular accents or dialects or make any other suitable selections. The combination sub-module 318 is configured to combine the created audio notification with the audio content 310. The created audio notification can be embedded into the audio content 310 at the identified insertion portion. The combination sub-module 318 provides the combined audio content 320 as an output. The combined audio content 320 comprises the original audio content 310 with the generated audio notification 322 inserted at the identified one or more insertion portions. This combined audio content 320 can be specific to the use case scenario. The combined audio content 320 is based on the content of the incoming communication 300 and also the original audio content 310. Therefore, the combined audio content would be different for different communication 300 and different audio content 310. In the example of Fig. 3 the combined audio content 320 can be played back to the user 106. The combined audio content 322 can be played back via a peripheral device 108 or any other suitable means. The combined audio content 320 enables the user 106 to receive the notification regarding the communication 300 without disrupting the audio content that they are listening to. The respective sub-modules 314, 316, 318 can be implemented using trained machine learning models and / or using any other suitable means. The machine learning models that are used can be stored in a memory of the apparatus 100 and / or can be accessed from any suitable location. The machine learning models can be a pre-trained machine learning model. The machine learning model can comprise a neural network or any other suitable type of trainable model. The term “Machine Learning Model” refers to any kind of artificial intelligence (Al), intelligent or other method that is trainable or tuneable using data. The machine learning model can comprise a computer program. The machine learning model can be trained to perform a task, such as identifying an insertion portion or creating an audio notification, without being explicitly programmed to perform that task. The machine learning model can be configured to learn from experience E with respect to some class of tasks T and performance measure P if its performance at tasks in T, as measured by P, improves with experience E. The machine learning model can learn from reference data to make estimations on future data. The machine learning model can be a trainable computer program. Other types of machine learning models could be used in other examples. The training of the machine learning model could be performed by a system that is separate to the apparatus 100. For example, the machine learning model could be trained by a system or other apparatus that has a higher processing capacity than the apparatus 100 of Fig. 1. In some examples the machine learning model could be trained by a system comprising one or more graphical processing units (GPUs) or any other suitable type of processor. The training data used to train the machine learning model can comprise stored audio content, stored communications and / or any other suitable data. The training of the machine learning models can be repeated as appropriate until the machine learning models have attained a sufficient level of stability. A machine learning model has a sufficient level of stability when fluctuations in the predictions provided by the machine learning model are low enough to enable the machine learning model to be used to identify an insertion portion or generate an audio notification or perform any other suitable function. The machine learning model has a sufficient level of stability when fluctuations in the predictions provided by the machine learning model are low enough so that the machine learning model provides consistent responses to test inputs. As an implementation example a user 106 can be listening to songs using a streaming application. In this example the user 106 is listening to Hotel California by the Eagles. While they are listening to this song the apparatus 100 obtains a communication 300. In this example the communication 300 is a text message. The text message can be received via any suitable messaging application. In this example the message is from the user’s friend Alice and comprises an invitation to go to her house for a meal the next evening. The apparatus 100 can use the examples of the disclosure to create a personalized audio notification and combine this with the song that the user 106 is listening to. In this case the content of the message from Alice is converted to lyrics that mimic the style and / or meter of the lyrics of Hotel California. The style could comprise the vocabulary used, or other characteristic, of the lyrics. For instance a first database of vocabulary could be used to generate lyrics for adding to country style songs while other databases of vocabulary could be used to generate lyrics for adding to different styles of songs such as rap or rock. These new lyrics are then inserted into an identified insertion portion. The identified insertion portion could be an instrumental portion or any other suitable portion. As an example, the Lyrics of the combined audio output 320 could be: Welcome to the Hotel California Such a lovely place (such a lovely place) Such a lovely face [A message arrives, from Alice, A friend with a heart, beyond compare, An invite to dine, tomorrow evening, In her home's warm glow, a magical space] Plenty of room at the Hotel California Any time of year (any time of year) You can find it here In this example the section in square brackets comprises the generated audio notification 322. The melody and voice used for this audio notification can be attuned or matched to the tune and voice used for the original song. As another implementation example a user 106 could be listening to audio content when the apparatus receives a communication 300. In this example the communication 300 is a voice note. The voice note could be received via any suitable messaging or communication application. In this example the voice note is from the user’s Auntie Susan asking for their power tools back. The apparatus 100 can use the examples of the disclosure to create a personalized audio notification and combine this with the audio content that the user 106 is listening to. To generate the personalized audio notification key information is extracted from the voice note and / or metadata of the voice note. For instance, a speech to text algorithm could be applied to the voice note to generate a text-based component and then an NLP module 306 or other suitable means can be used to extract the key information from the text-based component. The extracted information can then be used to create the personalized audio notification. In this case the personalized audio notification can be attuned to the audio content that the user is listening to and can indicate that Auntie Susan has messaged about her power tools. Fig. 4 shows an example controller. The controller 102 could be provided within any suitable apparatus 100, such as a telephone, a smart speaker, a computing device, a gaming device, a headset, or any other suitable type of apparatus. Implementation of the controller 102 may be as controller circuitry. The controller 102 may be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware). As illustrated in Fig. 4 the controller 102 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 406 in a general-purpose or special-purpose processor 402 that may be stored on a computer readable storage medium (disk, memory etc.) to be executed by such a processor 402. The processor 402 is configured to read from and write to the memory 404. The processor 402 may also comprise an output interface via which data and / or commands are output by the processor 402 and an input interface via which data and / or commands are input to the processor 402. The memory 404 stores a computer program 406 comprising computer program instructions (computer program code) that controls the operation of the controller 102 when loaded into the processor 402. The computer program instructions, of the computer program 406, provide the logic and routines that enables the apparatus 100 to perform the methods illustrated in the Figs. The processor 402 by reading the memory 404 is able to load and execute the computer program 406. The controller 102 therefore comprises: at least one processor 402; and at least one memory 404 storing instructions that, when executed by the at least one processor 402, cause an apparatus 100 at least to perform: identifying 200 one or more insertion portions of audio content 310 that is for playback to a user 106; obtaining 202 a communication 300; extracting 204 information from the obtained communication 300, wherein the information is extracted using language processing; generating 206 one or more audio notifications based on the extracted information wherein the one or more audio notifications are generated to be in a format that is attuned with the audio content 310; and adding 208 the one or more generated audio notifications to the identified one or more insertion portions of the audio content 310 to generate a combined output. The computer program 406 may arrive at the apparatus 100 via any suitable delivery mechanism 408. The delivery mechanism 408 may be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid-state memory, an article of manufacture that comprises or tangibly embodies the computer program 406. The delivery mechanism may be a signal configured to reliably transfer the computer program 406. The apparatus may propagate or transmit the computer program 406 as a computer data signal. The computer program 406 can comprise computer program instructions for causing an apparatus 100 to perform at least the following or for performing at least the following: identifying 200 one or more insertion portions of audio content 310 that is for playback to a user 106; obtaining 202 a communication 300; extracting 204 information from the obtained communication 300, wherein the information is extracted using language processing; generating 206 one or more audio notifications based on the extracted information wherein the one or more audio notifications are generated to be in a format that is attuned with the audio content 310; and adding 208 the one or more generated audio notifications to the identified one or more insertion portions of the audio content 310 to generate a combined output The computer program instructions may be comprised in a computer program, a non-transitory computer readable medium, a computer program product, a machine-readable medium. In some but not necessarily all examples, the computer program instructions may be distributed over more than one computer program. Although the memory 404 is illustrated as a single component / circuitry it may be implemented as one or more separate components / circuitry some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cached storage. Although the processor 402 is illustrated as a single component / circuitry it may be implemented as one or more separate components / circuitry some or all of which may be integrated / removable. The processor 402 may be a single core or multi-core processor. References to 'computer-readable storage medium’, ‘computer program product’, ‘tangibly embodied computer program’ etc. or a ‘controller’, ‘computer’, ‘processor’ etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc. As used in this application, the term ‘circuitry’ may refer to one or more or all of the following: (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g. firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device. The stages illustrated in the Figs, can represent steps in a method and / or sections of code in the computer program 406. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the block may be varied. Furthermore, it can be possible for some blocks to be omitted. The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to “comprising only one...” or by using “consisting”. In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected / coupled / in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e., so as to provide direct or indirect connection / coupling / communication. Any such intervening components can include hardware and / or software components. As used herein, the term "determine / determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also, "determine / determining" can include resolving, selecting, choosing, establishing, and the like. In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’ or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all of the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example. Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims. Features described in the preceding description may be used in combinations other than the combinations explicitly described above. Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not. Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not. The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a / an / the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning. The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and also to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result. In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described. The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure. Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance it should be understood that the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to and / or shown in the drawings whether or not emphasis has been placed thereon. l / we claim:
Claims
1. An apparatus comprising means for:identifying one or more insertion portions of audio content that is for playback to a user;obtaining a communication;extracting information from the obtained communication, wherein the information is extracted using language processing;generating one or more audio notifications based on the extracted information wherein the one or more audio notifications are generated to be in a format that is attuned with the audio content; andadding the one or more generated audio notifications to the identified one or more insertion portions of the audio content to generate a combined output.
2. An apparatus as claimed in claim 1 wherein identifying one or more insertion portions comprises determining a voice component within the audio content and identifying one or more portions where the voice component is below a threshold.
3. An apparatus as claimed in any preceding claim wherein identifying one or more insertion portions comprises determining one or more portions in the audio content where one or more parameters of the audio content are below a threshold.
4. An apparatus as claimed in any preceding claim wherein identifying one or more insertion portions comprises identifying one or more portions of the audio content that are replaceable with an audio notification.
5. An apparatus as claimed in any preceding claim wherein a machine learning model is used to identify the one or more insertion portions.
6. An apparatus as claimed in any preceding claim wherein the means are for adapting a duration of the one or more insertion portions to correspond to a duration of the generated audio notification.
7. An apparatus as claimed in any preceding claim wherein the obtained communication comprises, an incoming message, a notification from an application, or a calendar notification.
8. An apparatus as claimed in any preceding claim wherein the communication comprises text.
9. An apparatus as claimed in any preceding claim wherein a machine learning model is used to perform the language processing.
10. An apparatus as claimed in any preceding claim wherein attuning the audio notification with the audio content comprises adapting one or more parameters of the audio notification to correspond to the audio content wherein the one or more parameters comprise at least one of:tempo, pitch, language style, voice style11. An apparatus as claimed in any preceding claim wherein the means are for enabling the combined output to be played back to a user.
12. An apparatus as claimed in any preceding claim wherein the apparatus is at least one of: a telephone, a smart speaker, a computing device, a gaming device, or a headset.
13. A method comprising:identifying one or more insertion portions of audio content that is for playback to a user; obtaining a communication;extracting information from the obtained communication, wherein the information is extracted using language processing;generating one or more audio notifications based on the extracted information wherein the one or more audio notifications are generated to be in a format that is attuned with the audio content; andadding the one or more generated audio notifications to the identified one or more insertion portions of the audio content to generate a combined output.
14. A method as claimed in claim 13 wherein identifying one or more insertion portions comprises determining a voice component within the audio content and identifying one or more portions where the voice component is below a threshold.
15. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform at least:identifying one or more insertion portions of audio content that is for playback to a user;obtaining a communication;extracting information from the obtained communication, wherein the information is extracted using language processing;generating one or more audio notifications based on the extracted information wherein5 the one or more audio notifications are generated to be in a format that is attuned with the audio content; andadding the one or more generated audio notifications to the identified one or more insertion portions of the audio content to generate a combined output.
Citation Information
Patent Citations
Generating personalized audio programs from text content
US20140122079A1
Systems and methods for providing notifications within a media asset without breaking immersion
US20220044669A1
ViewUS20140122079A1onEspacenetopensinnewtab
ViewUS20220044669A1onEspacenetopensinnewtab