System and method for generating an audio presentation

CN115428476BActive Publication Date: 2026-08-28GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080100075.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-22
Publication Date
2026-08-28
Estimated Expiration
2040-04-22

Smart Images

  • Figure CN115428476B_ABST
    Figure CN115428476B_ABST
Patent Text Reader

Abstract

Systems and methods for generating an audio presentation are provided. A method can include obtaining data indicative of an acoustic environment of a user; obtaining data indicative of one or more events; generating, by an artificial intelligence system, an audio presentation for the user based at least in part on the data indicative of the one or more events and the data indicative of the acoustic environment of the user; and presenting the audio presentation to the user. The acoustic environment can include at least one of a first audio signal played on a computing system or a second audio signal associated with a surrounding environment of the user. The one or more events can include at least one of information communicated to the user by the computing system or at least a portion of the second audio signal associated with the surrounding environment of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to systems and methods for generating audio presentations. More specifically, this disclosure relates to devices, systems, and methods for incorporating event-related audio signals into a user's acoustic environment at a specific time using an artificial intelligence system. Background Technology

[0002] Personal computing devices such as smartphones have already provided the ability to listen to audio-based content on demand across a variety of platforms and applications. For example, a person can listen to music and movies stored locally on their smartphone; stream movies, music, TV shows, podcasts, and other content from numerous free and subscription-based services; access multimedia content available on the internet; and so on. Furthermore, advancements in wireless speaker technology have allowed users to listen to such audio content in a variety of environments.

[0003] However, in typical implementations, users have only a binary choice regarding whether to present audio information. For example, when listening to audio content in noise cancellation mode, all external signals, including audio information that the user prefers to hear, can be eliminated. Furthermore, when a user receives any type of notification, message, alert, etc., on their phone, the audio information associated with these events is usually presented upon receipt, frequently interrupting any other audio content being played for the user. Summary of the Invention

[0004] The aspects and advantages of this disclosure will be set forth in part in the description which follows, or may be apparent from the description, or may be learned by practice of embodiments of this disclosure.

[0005] One example aspect of this disclosure relates to a method for generating an audio presentation for a user. The method may include obtaining data indicative of the user's acoustic environment by a portable user device including one or more processors. The user's acoustic environment may include at least one of a first audio signal played on the portable user device or a second audio signal associated with the user's surrounding environment detected via one or more microphones that are part of or communicatively coupled to the portable user device. The method may also include obtaining data indicative of one or more events by the portable user device. The one or more events may include at least one of information to be conveyed to the user by the portable user device or at least a portion of the second audio signal associated with the user's surrounding environment. The method may further include generating an audio presentation for the user by an on-device artificial intelligence system based at least in part on the data indicative of the one or more events and the data indicative of the user's acoustic environment by the portable user device. Generating the audio presentation may include determining a specific time to incorporate a third audio signal associated with the one or more events into the acoustic environment. The method may further include presenting the audio presentation to the user by the portable user device.

[0006] Another example aspect of this disclosure relates to a method for generating an audio presentation for a user. The method may include obtaining data indicative of a user's acoustic environment by a computing system including one or more processors. The user's acoustic environment may include at least one of a first audio signal played on the computing system or a second audio signal associated with the user's surrounding environment. The method may also include obtaining data indicative of one or more events by the computing system. The one or more events may include at least one of information to be conveyed to the user by the computing system or at least a portion of the second audio signal associated with the user's surrounding environment. The method may further include generating an audio presentation for the user by an artificial intelligence system via the computing system, at least in part, based on the data indicative of the one or more events and the data indicative of the user's acoustic environment. The method may also include presenting the audio presentation to the user by the computing system. Generating the audio presentation by the artificial intelligence system may include the artificial intelligence system determining a specific time for incorporating a third audio signal associated with the one or more events into the acoustic environment.

[0007] Another example aspect of this disclosure relates to a method for training an artificial intelligence system. The artificial intelligence system may include one or more machine learning models. The artificial intelligence system may be configured to generate an audio presentation for a user by receiving data on one or more events and incorporating first audio signals associated with the one or more events into the user's acoustic environment. The method may include obtaining data, indicative of one or more prior events associated with the user, by a computing system including one or more processors. The data indicative of the one or more prior events may include semantic content of the one or more prior events. The method may also include obtaining data, indicative of a user response to the one or more prior events, by the computing system. The data indicative of the user response may include at least one of one or more prior user inputs in response to the one or more prior events, one or more prior user interactions with the computing system, or an intervention preference received in response to the one or more prior events. The method may also include training the artificial intelligence system, including one or more machine learning models, by the computing system to incorporate audio signals associated with one or more future events into the user's acoustic environment, at least in part based on the semantic content of the one or more prior events associated with the user and the data indicative of the user's response to the one or more events. The artificial intelligence system may be a local artificial intelligence system associated with the user.

[0008] Other aspects of this disclosure apply to various systems, apparatuses, non-transitory computer-readable media, machine-readable instructions, and electronic devices.

[0009] These and other features, aspects, and advantages of this disclosure will become more readily understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of the disclosure and, together with the specification, serve to explain the principles of the disclosure. Attached Figure Description

[0010] A complete and familiar description of this disclosure is set forth in the specification with reference to the accompanying drawings, wherein:

[0011] Figure 1A A block diagram is depicted of an example system for generating audio presentations for users via an artificial intelligence system, according to an example aspect of this disclosure;

[0012] Figure 1B A block diagram of an example computing device according to an example aspect of this disclosure is depicted;

[0013] Figure 1C A block diagram of an example computing device according to an example aspect of this disclosure is depicted;

[0014] Figure 2A A block diagram of an example artificial intelligence system according to an example aspect of this disclosure is depicted;

[0015] Figure 2B A block diagram of an example artificial intelligence system according to an example aspect of this disclosure is depicted;

[0016] Figure 2C A block diagram of an example artificial intelligence system according to an example aspect of this disclosure is depicted;

[0017] Figure 2D A block diagram of an example artificial intelligence system according to an example aspect of this disclosure is depicted;

[0018] Figure 2E A block diagram of an example artificial intelligence system according to an example aspect of this disclosure is depicted;

[0019] Figure 2F A block diagram of an example artificial intelligence system according to an example aspect of this disclosure is depicted;

[0020] Figure 3 A graphical representation of the acoustic environment of a user is depicted according to an example aspect of this disclosure;

[0021] Figure 4A A graphical representation of multiple events, including communications, is depicted according to an example aspect of this disclosure;

[0022] Figure 4B A graphical representation depicting an example summary of several events according to an example aspect of this disclosure;

[0023] Figure 5 A graphical representation of an example barge intervention tactic according to an example aspect of this disclosure is depicted;

[0024] Figure 6A A graphical representation of an example slip intervention strategy according to an example aspect of this disclosure is depicted;

[0025] Figure 6B A graphical representation of an exemplary sliding intervention strategy according to an exemplary aspect of this disclosure is depicted;

[0026] Figure 7 A graphical representation of an example filtering intervention strategy according to an example aspect of this disclosure is depicted;

[0027] Figure 8A A graphical representation of an example stretching intervention strategy according to an exemplary aspect of this disclosure is depicted;

[0028] Figure 8B A graphical representation of an example extension intervention strategy according to an exemplary aspect of this disclosure is depicted;

[0029] Figure 9A A graphical representation of an example loop intervention strategy according to an example aspect of this disclosure is depicted.

[0030] Figure 9B A graphical representation of an example cyclic intervention strategy according to an example aspect of this disclosure is depicted;

[0031] Figure 9C A graphical representation of an example cyclic intervention strategy according to an example aspect of this disclosure is depicted;

[0032] Figure 9D A graphical representation of an example cyclic intervention strategy according to an example aspect of this disclosure is depicted;

[0033] Figure 10 A graphical representation of an example mobile intervention strategy according to an example aspect of this disclosure is depicted;

[0034] Figure 11 A graphical representation of an example coverage intervention strategy according to an example aspect of this disclosure is depicted;

[0035] Figure 12A A graphical representation of an example ducking intervention strategy according to an exemplary aspect of this disclosure is depicted;

[0036] Figure 12B A graphical representation of an example evasion intervention strategy according to an exemplary aspect of this disclosure is depicted;

[0037] Figure 13 A graphical representation of an example interference (glitch) intervention strategy according to an example aspect of this disclosure is depicted;

[0038] Figure 14 An example method for generating audio rendering is described according to an example aspect of this disclosure;

[0039] Figure 15 An example method for generating audio rendering according to example aspects of this disclosure is described; and

[0040] Figure 16 An example training method according to example aspects of this disclosure is described. Detailed Implementation

[0041] Generally, this disclosure relates to devices, systems, and methods for generating audio presentations for users. For example, a computing device, such as a portable user device (e.g., a smartphone, wearable device, etc.), can acquire data indicating the user's acoustic environment. In some embodiments, the acoustic environment may include a first audio signal played on the computing device and / or a second audio signal associated with the user's surroundings. The second audio signal may be detected via one or more microphones of the computing device. The computing device may also acquire data indicating one or more events. One or more events may include at least a portion of information to be conveyed to the user by the computing system and / or the second audio signal associated with the surroundings. For example, in various embodiments, one or more events may include communications received by the computing device (e.g., text messages, SMS messages, voice messages, etc.), audio signals from the surroundings (e.g., announcements via a PA system), notifications from applications operating on the computing device (e.g., application logos, news updates, etc.), or prompts from applications operating on the computing device (e.g., turn-by-turn guidance from a navigation application). The computing system can then use an artificial intelligence (“AI”) system, such as an on-device AI system, to generate an audio presentation for the user, based at least in part on data indicating one or more events and data indicating the acoustic environment. For example, the AI ​​system can use one or more machine learning models to generate the audio presentation. The computing system can then present the audio presentation to the user. For example, in some implementations, the computing system can play the audio presentation for the user on a wearable speaker device (e.g., earbuds).

[0042] More specifically, the systems and methods of this disclosure can allow audible information to be provided to users as part of an immersive audio user interface, just as a graphical user interface visually provides information to users. For example, advancements in computing technology have allowed users to increasingly connect through various computing devices, such as personal user devices (e.g., smartphones, tablets, laptops, etc.) and wearable devices (e.g., smartwatches, earphones, smart glasses, etc.). Such computing devices allow information to be provided to users in real-time or near real-time. For example, applications operating on computing devices can allow real-time and near real-time communication (e.g., phone calls, text / SMS messages, video conferencing), notifications can quickly inform users of accessible information (e.g., email logos, social media post updates, news updates, etc.), and prompts can provide users with real-time instructions (e.g., step-by-step guidance, calendar reminders, etc.). However, in typical implementations, users may only have a binary option regarding whether to provide such information (e.g., provide all or nothing).

[0043] Furthermore, while advancements in wireless audio technology have allowed users to listen to audio content in various environments, such as when wearing wearable speaker devices (e.g., a pair of earbuds), whether or not to present audio information to the user is often a binary decision. For example, a user receiving one or more text messages may typically hear the audio associated with each message received, or none at all. Moreover, the audio associated with the text message is usually provided upon receipt, often interrupting any audio content being played for the user. Similarly, when a user listens to audio content in noise-cancellation mode, all external noise is typically canceled out. Therefore, some audio information that the user might want to hear (e.g., announcements on a PA system about the user's upcoming flight or someone speaking to the user) may be canceled out and never reach the user. Consequently, in order for the user to interact with their surroundings, the user may have to stop playing audio content, or in some cases, remove the wearable speaker device entirely.

[0044] However, the devices, systems, and methods of this disclosure can intelligently manage audio information for the user and present it to the user at appropriate times. For example, a computing system such as a portable user device can acquire data indicating the user's acoustic environment. For example, the acoustic environment may include audio signals played on the computing system (e.g., music, podcasts, audiobooks, etc.). The acoustic environment may also include audio signals associated with the user's surroundings. For example, one or more microphones of the portable user device can detect audio signals in the surrounding environment. In some embodiments, one or more microphones may be integrated into a wearable audio device (such as a pair of wireless earbuds).

[0045] The computing system can also acquire data indicating one or more events. For example, data indicating one or more events may include information to be conveyed by the computing system to a user and / or audio signals associated with the user's surrounding environment. For example, in some embodiments, one or more events may include communications received by the computing system to the user (e.g., text messages, SMS messages, voice messages, etc.). In some embodiments, one or more events may include external audio signals received by the computing system, such as audio signals associated with the surrounding environment (e.g., PA announcements, verbal communications, etc.). In some embodiments, one or more events may include notifications from applications operating on the computing system (e.g., application logos, news updates, social media updates, etc.). In some embodiments, one or more events may include prompts from applications operating on the computing system (e.g., calendar reminders, navigation prompts, telephone ringtones, etc.).

[0046] Data indicating one or more events and data indicating the acoustic environment can then be fed into an AI system, such as one locally stored on a computing system. For example, the AI ​​system may include one or more machine learning models (e.g., neural networks, etc.). The AI ​​system may generate an audio representation for the user based at least in part on the data indicating one or more events and the data indicating the acoustic environment. Generating the audio representation may include determining specific times to incorporate audio signals associated with one or more events into the acoustic environment.

[0047] The computing system can then present an audio presentation to the user. For example, in some embodiments, the computing system can be communicatively coupled to an associated peripheral device. The associated peripheral device can be, for example, a speaker device, such as an earphone device coupled to the computing system via Bluetooth or other wireless connections. In some embodiments, the associated peripheral device, such as a speaker device (e.g., a wearable earphone device), can also be configured to play an audio presentation to the user. For example, the computing device of the computing system can operate to communicate audio signals to the speaker device, such as via a Bluetooth connection, and upon receiving the audio signal, the speaker device can audibly play an audio presentation for the user.

[0048] In some implementations, an AI system can determine the specific time to incorporate audio signals associated with one or more events into an acoustic environment by identifying intermittent periods (e.g., gaps) within the acoustic environment. For example, an intermittent period can be a segment of the acoustic environment that corresponds to a relatively quiet period compared to the rest of the environment. For instance, for a user listening to a streaming music playlist, an intermittent period can correspond to a transition between consecutive songs. Similarly, for a user listening to an audiobook, an intermittent period can correspond to a period between chapters. For a user on a phone call, an intermittent period can correspond to a time after the user hangs up. For a user conversing with another person, an intermittent period can correspond to a break in the conversation.

[0049] In some implementations, pauses can be identified before audio content is played to the user. For example, playlists, audiobooks, and other audio content can be analyzed, and pauses can be identified, such as via a server computing device located remotely to the user's computing device. The data indicating the pauses can be stored by the server computing system and provided to the user's computing device.

[0050] In some implementations, intervals can be identified in real-time or near real-time. For example, one or more machine learning models can analyze audio content playing on a user's computing device and can analyze upcoming portions of the audio content (e.g., a 15-second window of upcoming audio content that will play in the near future). Similarly, one or more machine learning models can analyze audio signals in an acoustic environment to identify intervals in real-time or near real-time. In some implementations, the AI ​​system can select intervals as specific times at which audio signals associated with one or more events are incorporated into the acoustic environment.

[0051] In some implementations, an AI system may determine the urgency of one or more events based at least in part on at least one of the following: a user's geographic location, a source associated with one or more events, or the semantic content of data indicating one or more events. For example, a notification of a change in the meeting's location may be more urgent when a user is driving to a meeting than when the user has not yet left for the meeting. Similarly, a user may not want to be provided with certain information (e.g., text messages, etc.) while at work (e.g., at their workplace), but may want to receive such information when at home. The AI ​​system may use one or more machine learning models to analyze a user's geographic location and determine the urgency of one or more events based on that location.

[0052] Similarly, the sources associated with an event can be used to determine the urgency of one or more events. For example, a communication from a user's spouse may be more urgent than a notification from a news app. Likewise, a departure flight announcement broadcast through a PA system may be more urgent than a radio advertisement played in the user's acoustic environment. AI systems can use one or more machine learning models to identify the sources associated with one or more events and determine the urgency of one or more events based on those sources.

[0053] The semantic content of one or more events can also be used to determine the urgency of those events. For example, a text message from a user's spouse about their child being sick at school might be more urgent than a text message from the user's spouse asking the user to buy a gallon of milk on their way home. Similarly, a notification from a security app operating on a phone indicating a potential intrusion might be more urgent than a notification from the app's security panel about low battery. AI systems can use one or more machine learning models to analyze the semantic content of one or more events and determine their urgency based on that semantic content.

[0054] Furthermore, in some implementations, the AI ​​system can summarize the semantic content of one or more events. For example, a user might receive multiple group text messages where the group is deciding whether and where to have lunch. In some implementations, the AI ​​system can use a machine learning model to analyze the semantic content of the multiple text messages and generate a summary of the text messages. For example, this summary could include the location and time the group has chosen for their group lunch.

[0055] Similarly, in some implementations, a single event can be summarized. For example, a user might be waiting at an airport to board their flight. The boarding announcement for that flight can be issued via a PA system and can include information such as destination, flight number, departure time, and / or other information. The AI ​​system can generate a summary for the user, such as "Your flight is now boarding."

[0056] In some implementations, the AI ​​system can generate audio signals based at least in part on one or more events and incorporate those audio signals into the user's acoustic environment. For example, in some implementations, a text-to-speech (TTS) machine learning model can convert text information into audio signals and incorporate those audio signals into the user's acoustic environment. For instance, a summary of one or more events can be played for the user during intervals in the acoustic environment (e.g., at the end of a song).

[0057] In some implementations, the AI ​​system can determine not to incorporate audio signals associated with an event into the acoustic environment. For example, the AI ​​system can incorporate highly urgent events into the acoustic environment while ignoring (e.g., not incorporating) non-urgent events.

[0058] In some implementations, the AI ​​system can generate an audio presentation by eliminating at least a portion of the audio signal associated with the user's surrounding environment. For example, a user may be listening to music in noise cancellation mode. The AI ​​system can obtain audio signals from the user's surrounding environment, which may include ambient or background noise (e.g., car noise and horns, neighbor conversations, restaurant noise, etc.) and discrete audio signals, such as announcements via a PA system. In some implementations, the AI ​​system can eliminate the portion of the audio signal corresponding to ambient noise while playing music for the user. Furthermore, the AI ​​system can generate an audio signal associated with a PA announcement (e.g., a summary) and can incorporate that audio signal into the acoustic environment, as described herein.

[0059] In some implementations, the AI ​​system may use one or more intervention strategies to incorporate audio signals associated with one or more events into the acoustic environment. For example, the intervention strategy may be used to incorporate audio signals associated with one or more events at a specific time.

[0060] As an example, some audio signals associated with one or more events may be more urgent than others, such as highly urgent text messages or navigation prompts that indicate a user's direction at a specific time. In such cases, the AI ​​system can integrate the audio signals associated with one or more events into the acoustic environment as quickly as possible. For instance, the AI ​​system could use an "intrusion" intervention strategy where the audio signal being played for the user on the computing system is interrupted to make room for the audio signals associated with one or more events.

[0061] However, other intervention strategies can be used to present audio information to the user in a less intrusive manner. For example, in some implementations, a "filtering" intervention strategy can be used, where the audio signal played to the user is filtered (e.g., only certain frequencies of the audio signal are played) while the audio signal associated with one or more events is played. A "stretching" intervention strategy can maintain and repeat a portion of the audio signal played on the computing system while the audio signal associated with one or more events is played (e.g., maintaining the note of a song). A "looping" intervention strategy can select a portion of the audio signal played on the computing system and repeat that portion while the audio signal associated with one or more events is played (e.g., looping a 3-second audio clip). A "move" intervention strategy can change the perceived direction of the audio signal played on the computing system while the audio signal associated with one or more events is played (e.g., from left to right, from front to back, etc.). A "overlay" intervention strategy can overlay the audio signal associated with one or more events onto the audio signal played on the computing system (e.g., simultaneously). "Dodge" intervention strategies can reduce the volume of an audio signal played on a computing system when an audio signal associated with one or more events is played (e.g., making the first audio signal quieter). "Interference" intervention strategies can be used to create flaws in the audio signal played on the computing system. For example, interference intervention strategies can be used to provide contextual information to a user, such as informing the user when to turn (e.g., in response to navigation prompts) or marking distance markers (e.g., per mile) while the user is running. The intervention strategies described herein can be used to incorporate audio signals associated with one or more events into the user's acoustic environment.

[0062] In some implementations, the AI ​​system can generate audio presentations based at least in part on user input describing the listening environment. For example, a user can select a specific listening environment from a variety of options, and that specific listening environment can describe whether more or less audio information associated with one or more events should be conveyed to the user.

[0063] In some implementations, the AI ​​system can be trained, at least in part, based on previous user input describing intervention preferences. For example, a training dataset can be generated by receiving one or more user inputs in response to one or more events. For instance, when a user receives a text message, the AI ​​system can ask the user (e.g., via a graphical or audio user interface) whether they wish to be notified of similar text messages in the future. The AI ​​system can be trained to determine, for example, whether and / or when to present the user with audio information associated with similar future events, using factors such as the sender of the text message, the user's location, the semantic content of the text message, and the user's chosen listening environment preferences.

[0064] In some implementations, the AI ​​system can be trained, at least in part, based on one or more previous user interactions with the computing system in response to one or more prior events. For example, in addition to or as an alternative to user input specifically requesting one or more events, the AI ​​system can generate a training dataset based, at least in part, on whether and / or how a user responds to one or more events. As an example, a user's quick response to a text message can indicate that similar text messages should have a higher urgency level than text messages that are ignored, not responded to, or not responded to for an extended period of time.

[0065] Training datasets generated by AI systems can be used to train those systems. For example, one or more machine learning models of an AI system can be trained to respond to events that a user has previously responded to or events that the user has indicated as the preferred response. Training datasets can also be used to train local AI systems stored on a user's computing device.

[0066] In some implementations, the AI ​​system can generate one or more anonymized parameters based on a local AI system and provide these anonymized parameters to a server computing system. For example, the server computing system can use a joint learning method to train a global model using multiple anonymized parameters received from multiple users. The global model can be provided to each user and can be used, for example, to initialize the AI ​​system.

[0067] The systems and methods disclosed herein can provide numerous technical effects and benefits. For example, various embodiments of the disclosed technology can improve the efficiency of conveying audio information to a user. For instance, some embodiments can allow more information to be provided to a user without prolonging the total duration for which audio information is conveyed to the user.

[0068] Additionally or alternatively, certain implementations can reduce unnecessary user distraction, thereby enhancing user safety. For example, the devices, systems, and methods of this disclosure can allow audio information to be communicated to the user while the user is performing other tasks such as driving. Furthermore, in some implementations, the user's audio information can be filtered, summarized, and intelligently communicated to the user at an appropriate time based on the content and / or context of the audio information. This can improve the efficiency of communicating such information to the user and enhance the user experience.

[0069] Various embodiments of the devices, systems, and methods disclosed herein enable the wearing of head-mounted speaker devices (e.g., earbuds) without impairing a user's ability to operate effectively in the real world. For example, important real-world announcements can be communicated to the user at appropriate times, ensuring that the user's ability to effectively consume audio via the head-mounted speaker device is not adversely affected.

[0070] The systems and methods disclosed herein also provide improvements to computing technology. Specifically, a computing device, such as a personal user device, can acquire data indicative of the user's acoustic environment. The computing device can also acquire data indicative of one or more events. The computing device can generate an audio presentation for the user via an on-device AI system, based at least in part on the data indicative of one or more events and the data indicative of the user's acoustic environment. The computing device can then present the audio presentation to the user, such as via one or more wearable speaker devices.

[0071] Exemplary embodiments of this disclosure will now be discussed in more detail with reference to the accompanying drawings.

[0072] Figure 1 depicts an example system for generating audio presentations for a user according to exemplary aspects of this disclosure. System 100 may include a computing device 102 (e.g., a user / personal / mobile computing device such as a smartphone), a server computing system 130, and a peripheral device 150 (e.g., a speaker device). In some embodiments, computing device 102 may be a wearable computing device (e.g., a smartwatch, earphones, etc.). In some embodiments, peripheral device 150 may be a wearable device (e.g., earphones).

[0073] The computing device 102 may include one or more processors 111 and memory 112. The one or more processors 111 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 112 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. In some embodiments, the memory may include temporary memory, such as an audio buffer, for temporary storage of audio signals. The memory 112 may store data 114 and instructions 116, which may be executed by the processor 111 to cause the user computing device 102 to perform operations.

[0074] The computing device 102 may include one or more user interfaces 118. The user interface 118 can be used by a user to interact with the user computing device 102, such as providing user input, such as selecting a listening environment, responding to one or more events, etc.

[0075] The computing device 102 may also include one or more user input components 120 for receiving user input. For example, the user input component 120 may be a touch-sensitive component (e.g., a touch-sensitive display 118 or a touchpad) sensitive to the touch of a user input object (e.g., a finger or stylus). In some embodiments, the touch-sensitive component may be used to implement a virtual keyboard. Other example user input components 120 include one or more buttons, a conventional keyboard, or other components that a user may use to provide user input. The user input component 120 may allow the user to provide user input, such as via the user interface 120 or in response to information displayed in the user interface 120.

[0076] The computing device 102 may also include one or more displays 122. Displays 122 may be, for example, displays configured to display various information to a user via a user interface 118. In some embodiments, the one or more displays 122 may be touch-sensitive displays capable of receiving user input.

[0077] The computing device 102 may also include one or more microphones 124. The one or more microphones 124 may be any type of audio sensor and associated signal processing component configured to generate audio signals associated with the user's surrounding environment. For example, ambient audio, such as restaurant noise, passing vehicle noise, etc., may be received by the one or more microphones 124, which may generate audio signals based on the user's surrounding environment.

[0078] According to another aspect of this disclosure, computing device 102 may also include an artificial intelligence (AI) system 125, which includes one or more machine learning models 126. In some embodiments, machine learning models 126 may be operable to analyze a user's acoustic environment. For example, the acoustic environment may include audio signals played by computing device 102. For example, computing device 102 may be configured to play various media files, and the associated audio signals may be analyzed by one or more machine learning models 126, as disclosed herein. In some embodiments, the acoustic environment may include audio signals associated with the user's surrounding environment. For example, one or more microphones 124 may acquire and / or generate audio signals associated with the user's surrounding environment. One or more machine learning models 126 may be operable to analyze audio signals associated with the user's surrounding environment.

[0079] In some implementations, one or more machine learning models 126 may be operated to analyze data indicating one or more events. For example, the data indicating one or more events may include information to be conveyed to a user by computing device 102 and / or audio signals associated with the user's surrounding environment. For example, in some implementations, one or more events may include communications received by computing device 102 to the user (e.g., text messages, SMS messages, voice messages, etc.). In some implementations, one or more events may include external audio signals received by computing device 102, such as audio signals associated with the surrounding environment (e.g., PA announcements, verbal communications, etc.). In some implementations, one or more events may include notifications from applications operating on the computing device (e.g., application logos, news updates, social media updates, etc.). In some implementations, one or more events may include prompts from applications operating on computing device 102 (e.g., calendar reminders, navigation prompts, phone ringtones, etc.).

[0080] In some implementations, one or more machine learning models 126 may be, for example, neural networks (e.g., deep neural networks) or other multi-layered nonlinear models that output various information used by the artificial intelligence system. Further reference will follow below. Figures 2A to 2F The discussion covers example artificial intelligence system 125 and associated machine learning model 126 based on example aspects of this disclosure.

[0081] AI system 125 can be stored on a device (e.g., on computing device 102). For example, AI system 125 can be a local AI system 125.

[0082] The computing device 102 may also include a communication interface 128. The communication interface 128 may include any number of components for providing network communication (e.g., transceivers, antennas, controllers, cards, etc.). In some embodiments, the computing device 102 includes a first network interface operable for communicating using short-range wireless protocols such as Bluetooth and / or Bluetooth Low Energy, a second network interface operable for communicating using other wireless network protocols such as Wi-Fi, and / or a third network interface operable for communicating via GSM, CDMA, AMPS, 1G, 2G, 3G, 4G, 5G, LTE, GPRS, and / or other wireless cellular networks.

[0083] The computing device 102 may also include one or more speakers 129. The one or more speakers 129 may be configured, for example, to audibly play audio signals (e.g., generate sound waves including voice, speech, etc.) for a user to hear. For example, the artificial intelligence system 125 may generate an audio presentation for a user, and the one or more speakers 129 may present the audio presentation to the user.

[0084] Referring again to Figure 1, system 100 may further include server computing system 130. Server computing system 130 may include one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 may store data 136 and instructions 138 executed by processor 132 to cause server computing system 130 to perform operations.

[0085] In some embodiments, the server computing system 130 includes one or more server computing devices or is otherwise implemented by one or more server computing devices. Where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0086] In some implementations, server computing system 130 may store or include AI system 140, which may include one or more machine learning models 142. Further reference will be made below. Figures 2A to 2F The discussion covers example artificial intelligence system 140 and associated machine learning model 142 based on example aspects of this disclosure.

[0087] In some implementations, AI system 140 may be a cloud-based AI system, such as a personal cloud AI system 140 unique to a particular user. AI system 140 may be operable to generate audio presentations for users via cloud-based AI system 140.

[0088] Server computing system 130 and / or computing device 102 may include a model trainer 146 that uses various training or learning techniques, such as backpropagation of error, to train artificial intelligence systems 125 / 140 / 170. In some embodiments, performing backpropagation of error may include performing truncated backpropagation over time. Model trainer 146 may perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the trained model.

[0089] Specifically, model trainer 146 can train one or more machine learning models 126 / 142 / 172 based on a training data set 144. Training data 144 may include, for example, training datasets generated by AI systems 125 / 140 / 170. For example, as will be described in more detail herein, training data 144 may include data indicating one or more prior events and associated user input describing intervention preferences. In some implementations, training data 144 may include data indicating one or more prior events and data indicating one or more prior user interactions with computing device 102 in response to one or more prior events.

[0090] In some implementations, server computing device 130 may implement model trainer 146 to train new models or update versions of existing models on supplementary training data 144. As an example, model trainer 146 may receive anonymized parameters associated with local AI system 125 from one or more computing devices 102 and may use a joint learning approach to generate global AI system 140. In some implementations, global AI system 140 may be provided to multiple computing devices 102 to initialize local AI system 125 on multiple computing devices 102.

[0091] Server computing device 130 may periodically provide computing device 102 with one or more updated versions of AI system 140 and / or machine learning model 142. The updated AI system 140 and / or machine learning model 142 may be sent to user computing device 102 via network 180.

[0092] Model trainer 146 may include computer logic for providing the desired functionality. Model trainer 146 may be implemented using hardware, firmware, and / or software that controls a general-purpose processor. For example, in some embodiments, model trainer 146 includes a program file stored on a storage device, loaded into memory 112 / 134, and executed by one or more processors 111 / 132. In other embodiments, model trainer 146 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium such as a RAM hard disk or an optical or magnetic medium.

[0093] In some implementations, any process, operation, program, application, or instruction described as being stored at or executed by server computing device 130 may be wholly or partially stored at or executed by computing device 102, or vice versa. For example, as shown, computing device 102 may include a model trainer 146 configured to train one or more machine learning models 126 locally stored on computing device 102.

[0094] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication on network 180 can be carried over any type of wired and / or wireless connection using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0095] Referring again to Figure 1, system 100 may further include one or more peripheral devices 150. In some embodiments, peripheral device 150 may be a wearable speaker device, such as an earphone device that can be communicatively coupled to computing device 102.

[0096] Peripheral device 150 may include one or more user input components 152 configured to receive user input. User input components 152 may be configured to receive user interactions that request a response, such as in response to one or more events. For example, user input component 120 may be a touch-sensitive component (e.g., a touchpad) sensitive to the touch of a user input object (e.g., a finger or stylus). Other example user input components 152 include one or more buttons, switches, or other parts that a user can use to provide user input. User input components 152 may allow a user to provide user input such as requesting the display of one or more semantic entities.

[0097] Peripheral device 150 may also include one or more speakers 154. The one or more speakers 154 may be configured, for example, to audibly play audio signals (e.g., sound, speech, etc.) for a user to hear. For example, audio signals associated with media files played on computing device 102 may be communicated from computing device 102, such as via one or more networks 180, and these audio signals may be audibly played to the user by one or more speakers 154. Similarly, audio signals associated with communication signals received by computing device 102 (e.g., telephone calls) may be audibly played by one or more speakers 154.

[0098] Peripheral device 150 may also include a communication interface 156. Communication interface 156 may include any number of components providing network communication (e.g., transceivers, antennas, controllers, cards, etc.). In some embodiments, peripheral device 150 includes a first network interface operable for communication using short-range wireless protocols such as Bluetooth and / or Bluetooth Low Energy, a second network interface operable for communication using other wireless network protocols such as Wi-Fi, and / or a third network interface operable for communication via GSM, CDMA, AMPS, 1G, 2G, 3G, 4G, 5G, LTE, GPRS, and / or other wireless cellular networks.

[0099] Peripheral device 150 may also include one or more microphones 158. The one or more microphones 158 may be any type of audio sensor and associated signal processing component configured to generate audio signals associated with the user's surrounding environment. For example, ambient audio, such as restaurant noise, passing vehicle noise, etc., may be received by the one or more microphones 158, which may generate audio signals based on the user's surrounding environment.

[0100] Peripheral device 150 may include one or more processors 162 and memory 164. The one or more processors 162 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 164 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 164 may store data 166 and instructions 168 executed by processor 162 to cause peripheral device 150 to perform operations.

[0101] Peripheral device 150 may store or include AI system 170, which may include one or more machine learning models 172. Further references will follow. Figures 2A to 2FThis discussion covers an example artificial intelligence system 170 and an associated machine learning model 172 according to exemplary aspects of this disclosure. In some implementations, AI system 170 may be incorporated into or be part of AI systems 125 / 140. For example, AI systems 125 / 140 / 170 may be communicatively coupled and work together to generate audio presentations for a user. As an example, various machine learning models 124 / 142 / 172 may be stored locally on associated devices / systems 102 / 130 / 150 as part of AI systems 125 / 140 / 170, and machine learning models 124 / 142 / 172 may collectively generate audio presentations for a user.

[0102] For example, a first machine learning model 172 may acquire audio signals via a microphone 158 associated with the surrounding environment and perform noise cancellation on one or more portions of the audio signals acquired via the microphone 158. A second machine learning model 125 may incorporate event-associated audio signals into the noise-cancelled acoustic environment generated by the first machine learning model 172.

[0103] As described herein, AI system 170 may be trained by computing device 102 and / or server computing system 130 or otherwise provided to peripheral device 150.

[0104] Figure 1B A block diagram depicts an example computing device 10 implemented according to an exemplary embodiment of the present disclosure. The computing device 10 may be a user computing device or a server computing device.

[0105] Computing device 10 includes multiple applications (e.g., applications 1 to N). Each application contains its own machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.

[0106] like Figure 1B As shown, each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is application-specific.

[0107] Figure 1C A block diagram depicts an example computing device 50 implemented according to an example embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.

[0108] Computing device 50 includes multiple applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application may communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).

[0109] The central intelligence layer comprises multiple machine learning models. For example, such as... Figure 1C As shown, a corresponding machine learning model (e.g., a model) can be provided for each application and managed by a central intelligence layer. In other embodiments, two or more applications can share a single machine learning model. For example, in some embodiments, the central intelligence layer can provide a single model (e.g., a single model) for all applications. In some embodiments, the central intelligence layer is included in the operating system of computing device 50, or otherwise implemented by the operating system of computing device 50.

[0110] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data warehouse for computing device 50. For example... Figure 1C As shown, the central device data layer can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0111] Figure 2A A block diagram of an example AI system 200, comprising one or more machine learning models 202, is depicted according to exemplary aspects of this disclosure. In some embodiments, the AI ​​system 200 may be stored on a computing device / system, such as computing device 102, computing system 130, and / or peripheral device 150 depicted in FIG. 1. The AI ​​system 200 may be an AI system configured to generate an audio presentation 208 for a user. In some embodiments, the AI ​​system 200 is trained to receive data indicating one or more events 204.

[0112] For example, data indicating one or more events may include information to be conveyed to a user by the computing device / system and / or audio signals associated with the user's surrounding environment. For example, in some embodiments, one or more events may include communications received by the computing device / system to the user (e.g., text messages, SMS messages, voice messages, etc.). In some embodiments, one or more events may include external audio signals received by the computing device / system, such as audio signals associated with the surrounding environment (e.g., PA announcements, verbal communications, etc.). In some embodiments, one or more events may include notifications from applications operating on the computing device (e.g., application logos, news updates, social media updates, etc.). In some embodiments, one or more events may include prompts from applications operating on computing device 102 (e.g., calendar reminders, navigation prompts, telephone ringtones, etc.).

[0113] In some implementations, the AI ​​system 200 is trained to also receive data indicating the user's acoustic environment 206. For example, the data indicating the acoustic environment 206 may include audio signals (e.g., music, podcasts, audiobooks, etc.) played for the user on a computing device / system. The data indicating the acoustic environment 206 may also include audio signals associated with the user's surrounding environment.

[0114] like Figure 2A As shown, data indicating one or more events 204 and data indicating the acoustic environment 206 can be input into the AI ​​system 200, such as into one or more machine learning models 202. The AI ​​system 200 can generate an audio presentation 208 for the user based at least in part on the data indicating one or more events 204 and the data indicating the acoustic environment 206. For example, the audio presentation 208 (e.g., data indicating it) can be received as the output of the AI ​​system 200 and / or one or more machine learning models 202.

[0115] AI system 200 can generate audio presentation 208 by determining whether and when to incorporate audio signals associated with one or more events 204 into the acoustic environment 206. In other words, AI system 200 can intelligently manage audio information for the user.

[0116] For example, now refer to Figure 3 The figure depicts an example acoustic environment 300 for user 310. As shown, user 310 wears a wearable speaker device 312 (e.g., earbuds). In some embodiments, the acoustic environment 300 may include audio content played for user 310, such as music streamed from the user's personal computing device to the wearable speaker device 312.

[0117] However, the acoustic environment 300 of user 310 may also include additional audio signals, such as audio signals 320-328 associated with the user's surrounding environment. Each of audio signals 320-328 may be associated with a unique event. For example, as shown, audio signal 320 may be an audio signal generated by a musician on a loading platform at a train station. Another audio signal 322 may be an audio signal from the laughter of nearby children. Audio signal 324 may be an announcement via the PA system, such as an announcement that a particular train is boarding. Audio signal 326 may be an audio signal from a nearby passenger who shouts to attract the attention of other members of their travel group. Audio signal 328 may be an audio signal generated by a nearby train, such as an audio signal generated by a train traveling on the tracks or a horn indicating that a train is about to depart.

[0118] Harsh noise from the user's surroundings (audio signals 320-328) and any audio content played for the user (310) could be unbearable for the user (310). Therefore, in response, the user (310) who wishes to listen to audio content on their personal device can use noise cancellation mode to eliminate audio signals 320-328, thus allowing only audio content played on the user's personal device to be presented to the user. However, this may cause the user (310) to miss important audio information, such as announcements of the user's train's imminent departure broadcast via PA system 324. Therefore, in some cases, to ensure that the user (310) does not miss important audio content, the user (310) may have to turn off noise cancellation mode or completely remove the wearable speaker device 312.

[0119] Furthermore, even when user 310 is able to listen to audio content, such as audio content played on the user's personal device (e.g., a smartphone), such audio content may be frequently interrupted by other events, such as audio signals associated with communications, notifications, and / or prompts provided by the user's personal device. In response, the user can choose a "mute" mode, in which no audio signals associated with notifications on the device are provided, but this may also cause the user to similarly miss important information, such as text messages from a spouse or notifications from a travel app about travel delays.

[0120] See attached for reference. Figure 2A The AI ​​system 200 can intelligently manage a user's acoustic environment by determining whether and when to incorporate audio signals associated with one or more events into the user's acoustic environment. For example, according to an additional exemplary aspect of the invention, generating an audio presentation 208 by the AI ​​system 200 may include determining a specific time to incorporate audio signals associated with one or more events 204 into the acoustic environment 206.

[0121] For example, now refer to Figure 2B In some implementations, data indicating the acoustic environment 206 may be fed into one or more machine learning models 212 configured to identify intervals 214 within the acoustic environment 206. For example, interval 214 may be a portion of the acoustic environment 206 corresponding to a relatively quiet period, compared to other parts of the acoustic environment 206. For instance, for a user listening to a streaming music playlist, interval 214 may correspond to a transition period between consecutive songs. Similarly, for a user listening to an audiobook, interval 214 may correspond to a period between chapters. For a user on a phone call, interval 214 may correspond to a period after the user hangs up. For a user conversing with another person, interval 214 may correspond to an interruption in the conversation. (See reference...) Figure 6A and Figure 6B Example interval 214 is described in more detail.

[0122] In some implementations, the pause 214 can be identified before playing audio content to the user. For example, playlists, audiobooks, and other audio content can be analyzed by one or more machine learning models 212, and the pause 214 can be identified, for example, by a server computing device located remotely to the user's computing device. The data indicating the pause 214 can be stored by the server computing system and made available to the user's computing device.

[0123] In some implementations, the interval 214 can be identified in real-time or near real-time. For example, one or more machine learning models 212 can analyze audio content playing on a user's computing device and can analyze upcoming portions of the audio content (e.g., a 15-second window of upcoming audio content that will be played in the near future). Similarly, one or more machine learning models 212 can analyze audio signals in the acoustic environment 206 to identify the interval 214 in real-time or near real-time.

[0124] In some implementations, the AI ​​system 200 may select interval 214 as a specific time to incorporate audio signals associated with one or more events into the acoustic environment 206. For example, data indicating interval 214 and data indicating one or more events 204 may be input into a second machine learning model 216, which may generate an audio presentation 208 during interval 214 by incorporating audio signals associated with one or more events 204 into the acoustic environment 206.

[0125] In some implementations, one or more intervention strategies may be used to incorporate audio signals associated with one or more events 204 into the acoustic environment 206. (See reference...) Figures 5 to 13 Example intervention strategies based on example aspects of this disclosure are described in more detail.

[0126] Now for reference Figure 2C In some implementations, the AI ​​system 200 can generate audio signals 224 associated with one or more events. For example, data indicating one or more events 204 can be input into one or more machine learning models 222 configured to generate audio signals 224 associated with one or more events 204. For example, a text-to-speech (TTS) machine learning model 222 can convert text (e.g., a text message) associated with one or more events 204 into audio signals 224. Similarly, other machine learning models 222 can generate audio signals 224 associated with other events 204. For example, in some implementations, one or more machine learning models 222 can generate tonal audio signals 224 that can convey the context of one or more events 204. For example, different audio signals 224 can be generated for different navigation prompts, such as using a first tone to indicate a right turn and using a second tone to indicate a left turn.

[0127] Audio signal 224 (e.g., data indicating it) and acoustic environment 206 (e.g., data indicating it) can be input into one or more machine learning models 226, which can generate an audio presentation (e.g., data indicating it) 208 for the user. For example, as described herein, audio signal 224 can be incorporated into acoustic environment 206.

[0128] Now refer to the appendix Figure 2D In some implementations, the AI ​​system 200 may generate audio signals at least in part based on the semantic content 234 of one or more events 204. For example, data indicating one or more events 204 may be input into one or more machine learning models 232 configured to determine the semantic content 234 of one or more events 204. For example, the semantic content 234 of a notification may be determined by analyzing announcements via a PA system in the user's surrounding environment, such as by using a machine learning model 232 configured to convert speech to text. Furthermore, in some implementations, the semantic content 234 may be input into one or more machine learning models 236 configured to generate a summary 238 of the semantic content 234.

[0129] For example, the acoustic environment 206 of a user sitting at an airport may occasionally include PA system announcements with information about various flights, such as flight destination, flight number, departure time, and / or other information. However, the user may only want to hear announcements about his / her upcoming flight. In some implementations, the semantic content 234 of each flight announcement (e.g., each event) may be determined by one or more machine learning models 232. For most events 204 (e.g., most flight announcements), after analyzing the semantic content, the AI ​​system 200 may determine that the audio signal associated with event 204 does not need to be incorporated into the user's acoustic environment 206. For example, the AI ​​system 200 may determine not to incorporate audio signals associated with one or more events into the acoustic environment 204.

[0130] However, after obtaining the audio signal of the PA system announcement for a user's flight (e.g., a specific event), the AI ​​system 200 can determine that the audio signal associated with the announcement should be incorporated into the user's acoustic environment 206. For example, the AI ​​system 200 can recognize that the flight number in the semantic content 234 of the announcement corresponds to the flight number stored in the boarding pass document or calendar entry on the user's personal device.

[0131] In some implementations, AI system 200 can generate audio presentation 208 by selecting the current time period to provide the user with audio signals associated with one or more events. For example, AI system 200 can deliver a PA system notification about the user's flight to the user upon receipt, but noise-cancelled other notifications.

[0132] In some implementations, the AI ​​system can select future time periods to deliver audio signals associated with the announcement (e.g., during intervals, as described herein). However, while this approach can intelligently manage (e.g., filter) audio signals that the user may not care about, delivering or replaying PA announcements about a user's flight may present additional and unnecessary information beyond what the user needs.

[0133] To better manage the audio information presented to the user in audio presentation 208, in some implementations, the semantic content 234 of one or more events 204 can be summarized. For example, instead of replaying the PA system announcement to the user, one or more machine learning models 236 can use the semantic content 234 to generate a summary 238 of the announcement (e.g., a single event). For example, AI system 200 can generate summary 238 in which an audio signal with the information "Your flight is now boarding" is generated.

[0134] Similarly, in some implementations, multiple events can be summarized for the user. For example, now refer to Figure 4AThe text describes an example acoustic environment 410 for a user. The acoustic environment 410 may be, for example, audio content played for the user over a period of time. At different times, the user may receive text messages 420A-D (e.g., events) via the user's personal device. Each of the text messages 420A-D may be an event corresponding to an associated reception time 430A-D, as depicted with reference to the acoustic environment 410. For example, text messages 420A-D may be a chain of text messages of a group of people trying to decide whether and where to have lunch. Each of the events 420A-D (e.g., text messages 420A-D) may be input into an AI system, and corresponding semantic content may be determined for event 420A-D. Furthermore, now refer to... Figure 4B Summary 430 can be generated based at least in part on the semantic content of events 420A-D. For example, summary 430 can summarize the semantic content of text messages 420A-D, where summary 430 indicates that the group has decided to have tacos for lunch.

[0135] Although Figure 4A and 4B Various notifications and summaries are visually depicted, but information associated with the events and summaries can be provided to the user as audio content. For example, a summary 440 of text messages 420A-D can be incorporated into an acoustic environment 410 for playback to the user. For example, an audio signal 450 can be generated by an AI system 200, and the audio signal 440 can be incorporated into the acoustic environment 410. For example, as described herein, a text-to-speech machine learning model can audibly play summaries to the user during intervals (or other specific times) within the acoustic environment 410.

[0136] Now refer to the appendix Figure 2E In some implementations, the AI ​​system 200 may generate the audio presentation 208 based at least in part on the urgency 246 of one or more events. For example, as shown, in some implementations, the semantic content 234 of one or more events, the geographic location 240, and / or the source 242 associated with one or more events may be input into one or more machine learning models 244 to determine the urgency 246 of one or more events. The semantic content 234 may be, for example, semantic content generated by one or more machine learning models 232, such as… Figure 2D As shown.

[0137] For example, a user's geographic location 240 can indicate the user's acoustic environment and / or the user's preferences. For instance, when a user is at their workplace, they may prefer to be provided with audio content that is only associated with certain sources 242 and / or in which the semantic content 234 is particularly important and / or relevant to the user's work. However, when a user is at home, they may prefer to be provided with audio content associated with a wider and / or different set of sources 242, and / or in which the semantic content 234 is associated with a wider and / or different set of topics.

[0138] Similarly, when a user is moving, they may not want to be served certain audio content. For example, AI system 200 can use one or more machine learning models 246 to determine that the user is moving based on the user's changed geographic location 240 as they move. For instance, the user's changed geographic location 240 along a street could indicate that the user is driving. In this case, one or more machine learning models 244 can use the geographic location 240 to determine that only events with relatively high urgency 246 should be incorporated into the audio presentation 208.

[0139] As an example, a text message (e.g., semantic content 234) received by a user at their workplace (e.g., geographic location 240) from their spouse (e.g., source 242) indicating that the user's child is sick at school can be determined by one or more machine learning models 244 to have relatively high urgency 246. Conversely, a text message received by a user at their workplace (e.g., geographic location 240) from their spouse (e.g., source 242) requesting the user to buy a gallon of milk on their way home can be determined by one or more machine learning models 244 to have relatively low urgency 246.

[0140] Similarly, a user driving to the airport (e.g., location 240) receiving a text message (e.g., semantic content 234) from a friend (e.g., source 242) asking if the user wants to watch a baseball game can be determined by one or more machine learning models 244 to have relatively low urgency 246. In contrast, a notification received while the user is en route to the airport (e.g., location 240) from a travel app operating on the user's smartphone (e.g., source 242) informing the user that their upcoming flight has been delayed (e.g., semantic content 234) can be determined by one or more machine learning models 2442 to have relatively high urgency 246.

[0141] In some implementations, other data may also be used to determine urgency 246. For example, one or more contextual signifiers (not shown) may also be used to determine urgency 246. As an example, the time of day (e.g., during a user's typical workday) may indicate that the user is likely working, even if the user is at home (e.g., working remotely). Similarly, a particular day of the week (e.g., the weekend) may indicate that the user is likely not working. Furthermore, the activity the user is performing may also be a contextual signifier. As an example, a user editing a document or drafting an email may indicate that the user is performing a work activity. Similarly, a user navigating to a destination (e.g., driving a vehicle) may indicate that the user is busy and therefore should not be frequently interrupted. In this case, one or more machine learning models 248 may use such contextual signifiers to generate audio presentation 208.

[0142] The urgency 246 of event 204 and the user's acoustic environment 206 can be input into one or more machine learning models 248 to generate an audio presentation 208. For example, the urgency 246 of event 204 can be used to determine whether, when, and / or how audio signals associated with event 204 are incorporated into the acoustic environment 206. For example, an event 204 with a relatively high urgency 246 may be incorporated into the acoustic environment 206 faster than an event 204 with a relatively low urgency 246. Furthermore, different tones can be used to identify both the type of notification and its associated urgency. For example, a first frequency (e.g., low frequency) beeping tone may indicate that a low-urgency text message has been received, while a second frequency (e.g., high frequency) beeping tone may indicate that a high-urgency text message has been received. In this way, AI system 200 can generate audio presentation 208 by incorporating audio signals associated with one or more events 204 into the acoustic environment 206 based at least in part on the urgency 246 of one or more events 204.

[0143] Now for reference Figure 2F In some implementations, AI system 200 can generate an audio presentation by eliminating at least a portion of the audio signal associated with acoustic environment 206. For example, as depicted, acoustic environment 206 (e.g., data indicating it) can be fed into one or more machine learning models 252 to generate noise cancellation 254 (e.g., canceled audio signal). As an example, one or more machine learning models 252 of AI system 200 can perform active noise cancellation to allow certain ambient sounds (e.g., rain, bird chirping, etc.) to pass through while eliminating more jarring, disruptive sounds (car horns, shouting, etc.). Noise cancellation 254 can be incorporated into the audio presentation, such as audio presentation 208 described herein.

[0144] Normal reference Figures 2A to 2F The AI ​​system 200 and associated machine learning models can work together to intelligently manage the user's acoustic environment 206. For example, event 204 can be analyzed to determine the urgency 246 of event 204. Event 204 can be summarized based on its semantic content 234. The audio signal associated with event 204 can be generated by the AI ​​system 200. The AI ​​system 200 can determine the specific time to present the audio signal to the user, such as at a convenient time. The audio signal can be incorporated into the user's acoustic environment 206 at that specific time, such as music playing on the user's smartphone.

[0145] Furthermore, in some implementations, the AI ​​system may generate audio presentation 208 for the user based at least in part on user input describing the listening environment. For example, the user may select one of several different listening environments, which may include various thresholds for presenting audio information to the user. For example, at one end of the range, the user may select a real-time notification mode, in which each event with associated audio signals is presented to the user in real-time or near real-time. At the other end of the range, the user may select a mute mode, in which all external sounds in the surrounding environment are eliminated. One or more intermediate modes may include a summary mode in which events are summarized, an environment update mode in which white noise is generated and tonal audio information (e.g., indicating the tones of various events) is provided, and / or an environment mode in which only audio content from the user's surroundings is provided. When the user changes her listening mode, the AI ​​system 200 may adjust how audio information is incorporated into her acoustic environment 206.

[0146] According to additional examples of this disclosure, in some implementations, one or more intervention strategies may be used to incorporate audio signals associated with one or more events into the user's acoustic environment. Reference now is made to... Figure 5 This document describes an example "intrusion" intervention strategy. For example, an acoustic environment 510 is depicted, and the acoustic environment may include one or more audio signals, as described herein. In some implementations, the AI ​​system may use the intrusion strategy to interrupt the acoustic environment 510 to merge audio signals 520 associated with one or more events. For example, as shown, the audio signals of the acoustic environment 510 completely stop, while audio signals associated with one or more events 520 are played. Once the audio signals associated with one or more events 520 have been played, the acoustic environment 510 is restored. For example, the intrusion strategy may be used for events with relatively high urgency.

[0147] Now for reference Figure 6A and Figure 6B This describes an example of a "sliding" intervention strategy. For example, such as... Figure 6AAs shown, an acoustic environment 610 is illustrated. At 612, an intermittent occurs. For example, as described herein, intermittent 612 can correspond to a relatively quiet portion of the acoustic environment 610. Figure 6B As shown, by playing audio signal 620 during interval 612, audio signals associated with one or more events 620 can be incorporated into the acoustic environment 610. For example, a sliding intervention strategy can be used for events that do not have a relatively high degree of urgency, or to present audio information at a time that is more convenient or appropriate for the user.

[0148] Now for reference Figure 7 This describes an example of a "filtering" intervention strategy. For example, such as... Figure 7 As shown, an acoustic environment 710 is illustrated. At 712, a filtering strategy is applied to the acoustic environment 710. For example, as shown, only certain frequencies are allowed to pass. Audio signals associated with one or more events 720 can then be incorporated into the acoustic environment 710 by playing audio signal 720 when filtering 712 occurs.

[0149] Now for reference Figure 8A and Figure 8B This describes an example of a "stretching" intervention strategy. For example, such as... Figure 8A As shown, an acoustic environment 810 is illustrated, such as Figure 8B As shown, the acoustic environment 810 has been "stretched" by maintaining and continuously playing the first portion of the first audio signal. For example, the pitch of a song can be maintained for a period of time. When the acoustic environment 810 is stretched, audio signals associated with one or more events 820 can then be incorporated into the acoustic environment 810 by playing audio signal 820 when the stretch occurs.

[0150] Now for reference Figures 9A to 9D This illustrates an example of a "cyclic" intervention strategy. For example, as shown in the figure. Now refer to... Figure 9A An acoustic environment 910 is shown. A portion 912 (e.g., a segment) of the acoustic environment 910 can be selected. For example, portion 912 could be an upcoming portion of the acoustic environment 910, where audio signals associated with one or more events 920 will be incorporated into the acoustic environment 910. Figure 9B As shown, when section 912A is played (e.g., when the acoustic environment 910 reaches the first section), audio signals associated with one or more events 920 can be incorporated into the acoustic environment 910 by playing audio signal 920. Figure 9CAs shown, upon completion of playback portion 912A, the second portion 912B can be played simultaneously with the playback of audio signal 920. Upon completion of playback portion 912B, the third portion 912C can be played simultaneously with the playback of audio signal 920. The continuous portion 912 can be similarly repeated until the audio signal 920 is complete. In this way, the loop intervention strategy can maintain and repeatedly play portions 912 of the acoustic environment 910 by repeatedly looping portions 912 of the acoustic environment 910.

[0151] Now for reference Figure 10 This describes an example of a "mobility" intervention strategy. For example, such as... Figure 10 As shown, an acoustic environment 1010 is illustrated. As shown, the perceived direction of the acoustic environment 1010 can be changed when an audio signal associated with one or more events 1020 is played. For example, the perceived direction of the acoustic environment 1010 can be changed by moving the stereo acoustic environment 1010 from the left to the right, from the front to the back, etc. In some embodiments, changing the perceived direction may include incorporating a "silencing" effect, wherein the acoustic environment 1010 is perceived as being at a distance from the user.

[0152] Now for reference Figure 11 This describes an example of a "coverage" intervention strategy. For example, such as... Figure 11 As shown, an acoustic environment 1110 is illustrated. As illustrated, by simultaneously playing both the acoustic environment 1110 and the audio signal 1120, the audio signal associated with one or more events 1120 is overlaid by the acoustic environment 1110. This overlay intervention strategy can be used to provide context to the user. For example, a first tone can be used to instruct the driver to turn left, while a second tone can be used to instruct a right turn.

[0153] Now for reference Figure 12A and Figure 12B This describes an example of an "avoidance" intervention strategy. For example, such as... Figure 12A As shown, an acoustic environment 1210 is illustrated. However, as... Figure 12B As shown, the volume of the acoustic environment 1210 has been reduced while playing an audio signal associated with one or more events 1220. A dodging intervention strategy can be used to gradually or abruptly reduce the volume of the acoustic environment 1210. The rate at which the volume of the acoustic environment 1210 is reduced can be used, for example, to provide context to the audio signal 1220, such as indicating the urgency of one or more events.

[0154] Now for reference Figure 13 This describes example "interference" intervention strategies. For example, such as... Figure 13As shown, an acoustic environment 1310 is illustrated. As illustrated, an audio signal associated with one or more events 1320 can be generated by creating defects in the acoustic environment 1310. For example, the defect could resemble a record scratch or a jump in a digital audio track. This defect can be used to provide context to a user. For example, a distraction strategy could be used by a runner listening to music to mark distances or time markers (e.g., per mile, per minute, etc.).

[0155] General reference Figures 5 to 13 As shown, the intervention strategies described herein can be used individually or in combination. For example, stretching and dodging strategies can be used to stretch and reduce the volume of the acoustic environment. Furthermore, it should be noted that the acoustic environment described herein may include the playback of audio content for the user and / or the cancellation of audio signals. For example, a user listening in ambient mode may allow certain sounds (e.g., rain) to be delivered to the user while other sounds (car horns) are canceled.

[0156] Figure 14 A flowchart is depicted for example method 1400 used to generate audio rendering. Although Figure 14 The steps are described in a specific order for illustrative and discussion purposes, but the method of this disclosure is not limited to the specific order or arrangement shown. The steps of method 1400 may be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of this disclosure.

[0157] At 1402, the method may include acquiring data indicating the acoustic environment. For example, in some embodiments, the data indicating the acoustic environment may include audio signals played for the user on the user's portable user device. In some embodiments, the data indicating the acoustic environment may include audio signals associated with the user's surrounding environment. For example, one or more microphones may detect / acquire audio signals associated with the surrounding environment.

[0158] At 1404, the method may include obtaining data indicating one or more events. For example, in some embodiments, the data indicating one or more events may be obtained by a portable user device. One or more events may include information such as that conveyed to a user by the portable user device, and / or a portion of an audio signal associated with the user's surrounding environment. In some embodiments, one or more events may include communications received by the portable user device to the user (e.g., text messages, SMS messages, voice messages, etc.). In some embodiments, one or more events may include external audio signals received by the portable user device, such as audio signals associated with the surrounding environment (e.g., PA announcements, verbal communications, etc.). In some embodiments, one or more events may include notifications from applications operating on the portable user device (e.g., application logos, news updates, social media updates, etc.). In some embodiments, one or more events may include prompts from applications operating on the portable user device (e.g., calendar reminders, navigation prompts, telephone ringtones, etc.).

[0159] In 1406, the method may include generating an audio presentation for a user by an AI system based at least in part on data indicating one or more events and data indicating the user's acoustic environment. For example, in some embodiments, the AI ​​system may be an on-device AI system for a portable user device.

[0160] At 1408, the method may include presenting an audio presentation to a user. For example, in some embodiments, the audio presentation may be presented by a portable user device. For example, the portable user device may present the audio presentation to the user via one or more wearable speaker devices, such as one or more earbuds.

[0161] Now for reference Figure 15 The flowchart depicts an example method 1500 for generating audio presentations for users. Although... Figure 15 The steps are described in a specific order for illustrative and discussion purposes, but the method of this disclosure is not limited to the specific order or arrangement shown. The various steps of method 1500 may be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of this disclosure.

[0162] In 1502, the method may include determining the urgency of one or more events. For example, in some implementations, the AI ​​system may use one or more machine learning models to determine the urgency of one or more events based at least in part on the user's geographic location, the source associated with one or more events, and / or the semantic content of one or more events.

[0163] In 1504, the method may include identifying intervals within an acoustic environment. For example, an interval may be part of the acoustic environment that corresponds to a relatively quiet period, compared to other parts of the environment. For instance, for a user listening to a streaming music playlist, an interval may correspond to a transition period between consecutive songs. Similarly, for a user listening to an audiobook, an interval may correspond to a time period between chapters. For a user on a telephone call, an interval may correspond to a period of time after the user hangs up. For a user conversing with another person, an interval may correspond to an interruption in the conversation.

[0164] In 1506, the method may include determining a specific time to incorporate audio signals associated with one or more events into an acoustic environment. For example, in some embodiments, the specific time may be determined (e.g., selected) based at least in part on the urgency of one or more events. For example, an event with relatively higher urgency may be presented earlier than an event with relatively lower urgency. In some embodiments, the AI ​​system may select an identified interval as the specific time to incorporate audio signals associated with one or more events. In some embodiments, determining the specific time to incorporate audio signals associated with one or more events may include determining not to incorporate audio signals into the acoustic environment. In some embodiments, determining the specific time may include determining a specific time to incorporate a first audio signal into the acoustic environment while determining not to incorporate a second audio signal.

[0165] In 1508, the method may include generating an audio signal. For example, in some embodiments, the audio signal may be a tone indicating the urgency of one or more events. In some embodiments, the audio signal associated with one or more events may include a summary of the semantic content of one or more events. For example, in some embodiments, an audio signal such as a summary may be generated by a text-to-speech (TTS) model.

[0166] In 1510, the method may include noise cancellation. For example, in some implementations, generating an audio representation for a user may include canceling one or more audio signals associated with the user's surrounding environment.

[0167] In 1512, the method may include incorporating audio signals associated with one or more events into the user's acoustic environment. For example, in some implementations, one or more intervention strategies may be used. For instance, the AI ​​system may use an intrusion intervention strategy whereby the audio signal being played for the user on the computing system is interrupted to make room for the audio signal associated with one or more events. In some implementations, the AI ​​system may use a sliding intervention strategy to play the audio signal associated with one or more events during intervals in the acoustic environment. In some implementations, a filtering intervention strategy may be used whereby the audio signal played for the user is filtered while playing the audio signal associated with one or more events (e.g., only certain frequencies of the audio signal are played). In some implementations, a stretching intervention strategy may be used whereby the AI ​​system maintains and repeats a portion of the audio signal played on the device while playing the audio signal associated with one or more events (e.g., maintaining the pitch of a song). In some implementations, a looping intervention strategy may be used whereby the AI ​​system selects a portion of the audio signal played on the device and repeats that portion while playing the audio signal associated with one or more events (e.g., looping a 3-second audio clip). In some implementations, a motion intervention strategy may be used, wherein the AI ​​system alters the perceived direction of the audio signal played on the computing system (e.g., from left to right, from front to back, etc.) while playing an audio signal associated with one or more events. In some implementations, an overlay intervention strategy may be used, wherein the AI ​​system overlays the audio signal played on the device with the audio signal associated with one or more events (e.g., simultaneously). In some implementations, a dodging intervention strategy may be used, wherein the AI ​​system reduces the volume of the audio signal played on the device (e.g., makes the first audio signal quieter) while playing an audio signal associated with one or more events. In some implementations, an interference intervention strategy may be used, wherein the AI ​​system generates defects in the audio signal played on the device.

[0168] Now for reference Figure 16 The flowchart depicts an example method 1600 for training an AI system. Although Figure 16 The steps are described in a specific order for illustrative and discussion purposes, but the method of this disclosure is not limited to the specific order or arrangement shown. The steps of method 1600 may be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of this disclosure.

[0169] In 1602, the method may include obtaining data indicating one or more prior events. For example, one or more prior events may include communications received by the computing system to a user (e.g., text messages, SMS messages, voice messages, etc.). In some embodiments, one or more events may include external audio signals received by the computing system, such as audio signals associated with the surrounding environment (e.g., PA announcements, verbal communications, etc.). In some embodiments, one or more events may include notifications from applications operating on the computing system (e.g., application logos, news updates, social media updates, etc.). In some embodiments, one or more events may include prompts from applications operating on the computing system (e.g., calendar reminders, navigation prompts, telephone rings, etc.). In some embodiments, the data indicating one or more prior events may be included in a training dataset generated by an AI system.

[0170] In 1604, the method may include obtaining data indicating a user's response to one or more prior events. For example, the data indicating a user's response may include one or more previous user interactions with the computing system in response to one or more prior events. For example, whether a user viewed a news article from a news app notification could be used to train whether similar news updates should be provided in the future. In some implementations, the data indicating a user's response may include one or more previous user inputs describing intervention preferences received in response to one or more prior events. For example, the AI ​​system may ask the user if they would like to receive similar content in the future. In some implementations, the data indicating a user's response may be included in a training dataset generated by the AI ​​system.

[0171] In 1606, the method may include training an AI system comprising one or more machine learning models to incorporate audio signals associated with one or more future events into the user's acoustic environment, based at least in part on the semantic content of one or more previous events associated with the user and data indicating the user's response to one or more events. For example, the AI ​​system may be trained to incorporate audio signals into the acoustic environment in a manner similar to how the user would react to similar events, or better still, in accordance with the user's stated preferences.

[0172] In 1608, the method may include determining one or more anonymized parameters associated with the AI ​​system. For example, the AI ​​system may be a local AI system stored on a user's personal device. The one or more anonymized parameters may include, for example, one or more anonymized parameters of one or more machine learning models of the AI ​​system.

[0173] In 1610, the method may include providing a server computing system with one or more anonymized parameters associated with an AI system, the server computing system being configured to determine a global AI system at least in part based on the one or more anonymized parameters via joint learning. For example, the server computing system may receive anonymized parameters for multiple local AI systems and may generate a global AI system. For example, the global AI system may be used to initialize an AI system on a user device.

[0174] This article discusses technologies related to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and received from these systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and partitions of tasks and functions among components. For example, the server processes discussed in this article can be implemented using a single server or multiple servers working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0175] While the subject matter has been described in detail with reference to specific exemplary embodiments and methods, it should be understood that those skilled in the art can readily generate modifications, variations, and equivalents of these embodiments based on the foregoing. Therefore, the scope of this disclosure is exemplary and not restrictive, and this disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the subject matter, which will be apparent to those skilled in the art.

[0176] Furthermore, although this disclosure is generally discussed with reference to computing devices such as smartphones, it is also applicable to other forms of computing devices, including, for example, laptop computing devices, tablet computing devices, wearable computing devices, desktop computing devices, mobile computing devices, or other computing devices.

Claims

1. A method for generating audio presentations for a user, comprising: Data indicative of a user's acoustic environment is obtained by a portable user equipment including one or more processors, the user's acoustic environment including at least one of a first audio signal played on the portable user equipment or a second audio signal associated with the user's surrounding environment detected via one or more microphones, the one or more microphones being part of or communicatively coupled to the portable user equipment; Data indicating one or more events is obtained from a portable user device, the one or more events including at least one portion of information to be conveyed by the portable user device to the user or a second audio signal associated with the user's surrounding environment; An on-device artificial intelligence system in a portable user device generates an audio presentation for a user based at least in part on data indicating one or more events and data indicating the user's acoustic environment. Generating the audio presentation includes: generating a third audio signal associated with one or more events by the artificial intelligence system based at least in part on semantic content of the data indicating one or more events; determining a specific time to incorporate the third audio signal into the acoustic environment; wherein the audio presentation includes a summary of one or more events based at least in part on the semantic content of the one or more events; and The audio presentation is delivered to the user via a portable user device.

2. The method according to claim 1, wherein, The audio is presented to the user via one or more wearable speaker devices.

3. The method according to claim 1, wherein, The first audio signal is played to the user via one or more headphone speaker devices and / or at least one of one or more microphones that form part of the one or more headphone speaker devices.

4. The method according to any of the preceding claims, wherein, The one or more wearable speaker devices include one or more head-mounted wearable speaker devices.

5. A method for generating audio presentations for a user, comprising: Data indicative of a user's acoustic environment is obtained by a computing system comprising one or more processors, the user's acoustic environment including at least one of a first audio signal played on the computing system or a second audio signal associated with the user's surrounding environment; The computing system obtains data indicating one or more events, the one or more events including at least one portion of information to be conveyed by the computing system to the user or a second audio signal associated with the user's surrounding environment; An artificial intelligence system generates an audio presentation for a user, at least in part, based on data indicating one or more events and data indicating the user's acoustic environment, via a computing system. The audio presentation includes a summary of one or more events, at least in part, based on the semantic content of the events. The audio presentation is then presented to the user by the computing system. The generation of the audio presentation by the artificial intelligence system includes: generating a third audio signal associated with one or more events by the artificial intelligence system based at least in part on the semantic content of data indicating one or more events; and determining a specific time for the artificial intelligence system to incorporate the third audio signal into the acoustic environment.

6. The method according to claim 5, wherein, The specific time at which the artificial intelligence system determines the incorporation of a third audio signal associated with one or more events into the acoustic environment includes: Identifying intermittent sounds in an acoustic environment; and Select the interval as the specific time.

7. The method according to claim 5, wherein, The specific time at which the artificial intelligence system determines the incorporation of a third audio signal associated with one or more events into the acoustic environment includes: The urgency of one or more events is determined by the artificial intelligence system based at least in part on at least one of the following: the user's geographic location, a source associated with one or more events, or the semantic content of data indicating the one or more events. The specific time is determined by the artificial intelligence system based at least in part on the urgency of one or more events.

8. The method according to any one of claims 5-7, wherein, A third audio signal is associated with a first event in one or more events, and wherein the method further includes: The artificial intelligence system determines that audio signals associated with a second event in one or more events will not be incorporated into the acoustic environment.

9. The method according to any one of claims 5-7, wherein, Obtaining data indicating the user's acoustic environment includes obtaining a second audio signal associated with the user's surroundings; and The audio presentation generated for the user by the artificial intelligence system includes noise cancellation of at least a portion of a second audio signal associated with the user's surrounding environment.

10. The method according to any one of claims 5-7, wherein, The audio presentation generated by the artificial intelligence system also includes the artificial intelligence system incorporating a third audio signal into the acoustic environment at the specific time.

11. The method according to any one of claims 5-7, wherein, Generating a third audio signal by the artificial intelligence system based at least in part on the semantic content of one or more events includes summarizing the semantic content of one or more events.

12. The method according to any one of claims 5-7, wherein, The one or more events include at least one of the following: a communication received by the computing system to the user; an external audio signal received by the computing system that includes at least a portion of a second audio signal associated with the user's surrounding environment; a notification from an application operating on the computing system; or a prompt from an application operating on the computing system.

13. The method according to any one of claims 5-7, wherein, Incorporating a third audio signal into an acoustic environment includes at least one of the following: incorporating a third audio signal into an acoustic environment using at least one intervention strategy; and The at least one intervention strategy includes at least one of the following: interrupting the first audio signal, filtering the first audio signal, maintaining and continuously playing the first part of the first audio signal by stretching the first part of the first audio signal, maintaining and repeatedly playing the second part of the first audio signal by repeatedly looping the second part of the first audio signal, changing the perception direction of the first audio signal, superimposing a third audio signal onto the first audio signal, reducing the volume of the first audio signal, or generating a defect in the first audio signal.

14. The method according to any one of claims 5-7, wherein, The determination by the artificial intelligence system of a specific time to incorporate a third audio signal associated with one or more events into the acoustic environment includes: the determination by the artificial intelligence system not to incorporate the third audio signal into the acoustic environment.

15. The method according to any one of claims 5-7, wherein, The audio presentation is generated at least in part based on user input describing the listening environment.

16. The method according to any one of claims 5-7, wherein, The artificial intelligence system has been trained, at least in part, based on previous user input describing intervention preferences.

17. The method according to any one of claims 5-7, wherein, The artificial intelligence system has been trained, at least in part, based on one or more previous user interactions with the computing system in response to one or more previous events.

18. A system for generating audio presentations for a user, comprising: Artificial intelligence systems that include one or more machine learning models; One or more processors; as well as One or more non-transitory computer-readable media, which collectively store instructions that, when run by one or more processors, cause a computing system to perform operations, said operations including: Obtain data indicating the user's acoustic environment, the user's acoustic environment including at least one of a first audio signal played on the computing system or a second audio signal associated with the user's surrounding environment; Obtain data indicating one or more events, said one or more events including at least a portion of information to be conveyed to a user by the computing system or a second audio signal associated with the user's surrounding environment; The artificial intelligence system generates an audio presentation for a user based at least in part on data indicating one or more events and data indicating the user's acoustic environment, wherein the audio presentation includes a summary of one or more events based at least in part on the semantic content of one or more events; and Presenting audio to the user; The audio generated by the artificial intelligence system includes: The artificial intelligence system generates a third audio signal associated with one or more events, based at least in part on the semantic content of data indicating one or more events; Determine the specific time to incorporate the third audio signal into the acoustic environment; and The third audio signal is incorporated into the acoustic environment at the specified time.

19. The system according to claim 18, wherein, The system also includes a wearable device containing a speaker; and Among them, presenting audio to users includes playing audio via wearable devices.

20. A portable user equipment comprising one or more processors configured via machine-readable instructions to perform the method of any one of claims 1 to 17.

21. A machine-readable instruction, when executed, causes the method of any one of claims 1 to 17 to be performed.

Citation Information

Patent Citations

  • Method and device for embedding event notification into multimedia content

    CN101238711A

  • Methods and apparatus for delivering ancillary information to the user of a portable audio device

    US20080037718A1

  • Methods and Apparatus to Assist Listeners in Distinguishing Between Electronically Generated Binaural Sound and Physical Environment Sound

    US20170359467A1