Method and system for integrating short-term context for content playback adaptation

CN116671113BActive Publication Date: 2026-09-18GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180078867.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-24
Filing Date
2021-11-11
Publication Date
2026-09-18
Estimated Expiration
2041-11-11

AI Technical Summary

Technical Problem

在其他情况下,来自设备的大声回放音频可能使得用户难以注意到需要用户注意的正在进行的事件,诸如定时器关闭、呼入电话呼叫或通过婴儿监视器的婴儿哭泣

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116671113B_ABST
    Figure CN116671113B_ABST
Patent Text Reader

Abstract

While an assistant-enabled device (10) is playing back media content (120), a method (400) includes receiving a contextual signal (102) from an environment of the assistant-enabled device and executing an event identification routine (200) to determine whether the received contextual signal indicates an event that conflicts with playback of the media content from the assistant-enabled device. When the event identification routine determines that the received contextual signal indicates an event that conflicts with playback of the media content, the method further includes adjusting a content playback setting of the assistant-enabled device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to integrating short-term scenarios for content playback adaptation. Background Technology

[0002] Using digital assistants to stream music from smart speakers and mobile devices is common in user environments such as homes or offices. Besides music, digital assistants are also used to play video content via smart players. Playback content from these devices (such as audio and / or video playback) can interfere with ongoing conversations or activities in the user's environment. For example, playback content might interfere with a conversation between two users in the environment or a conversation being conducted by a user on the phone. In such cases, the user will manually tune the device to control the playback content so that it does not interfere with the current user activity. For example, a user can walk to the smart speaker playing music and lower / mute the volume so that it no longer interferes with the user's activity. In other cases, loud playback audio from the device can make it difficult for the user to notice ongoing events that require their attention, such as a timer going off, an incoming phone call, or a baby crying through a baby monitor. Summary of the Invention

[0003] One aspect of this disclosure provides a method for adjusting playback settings of an assistant-enabled device. While the assistant-enabled device is playing media content, the method includes: receiving a context signal from the environment of the assistant-enabled device at the data processing hardware of the assistant-enabled device; executing an event recognition routine by the data processing hardware to determine whether the received context signal indicates an event that conflicts with the playback of media content from the assistant-enabled device; and adjusting the content playback settings of the assistant-enabled device by the data processing hardware when the event recognition routine determines that the received context signal indicates an event that conflicts with the playback of media content.

[0004] Embodiments of this disclosure include one or more of the following optional features. In some embodiments, the context signal includes at least one of the following: audio detected by the microphone of the assistant-enabled device; or image data captured by the image capture device of the assistant-enabled device. In other embodiments, the context signal includes network-based information from a user account shared with nearby devices in the environment of the assistant-enabled device. Here, the network-based information indicates an event associated with a nearby device. The context signal may include communication signals transmitted from nearby devices communicating with the assistant-enabled device. The communication signals indicate an event associated with a nearby device.

[0005] In some examples, performing an event recognition routine includes executing a neural network-based classification model configured to receive contextual signals as input and generate a classification result as output, indicating whether the received contextual signals indicate one or more events that conflict with playback of media content from an assistant-enabled device. In these examples, the contextual signals received as input to the neural network-based classification model include an audio stream, and the classification result generated by the neural network-based classification model as output includes audio events that conflict with playback of the media content. Furthermore, the classification result generated by the neural network-based classification model as output may be further based on the audibility level of the audio stream. Alternatively, in these examples, the contextual signals received as input to the neural network-based classification model include an image stream, and the classification result generated by the neural network-based classification model as output includes activity events that conflict with playback of the media content.

[0006] In some implementations, when the received context signal indicates an audio event, the method further includes: obtaining, by data processing hardware, an audible level associated with the audio event; obtaining, by data processing hardware, an audible level of media content played back from the assistant-enabled device; and determining, by data processing hardware, a probability score indicating the likelihood that the media content played back from the assistant-enabled device interrupts the ability of a user associated with the assistant-enabled device to hear the audio event. Here, adjusting the content playback settings of the assistant-enabled device includes, based on the probability score, performing one of the following: reducing the audible level of the media content played back from the assistant-enabled device; or stopping / pausing the playback of media content from the assistant-enabled device.

[0007] In another implementation, when the event recognition routine determines that the received context signal indicates an event that conflicts with the playback of media content, the method further includes: obtaining playback features associated with the media content being played back from the assistant-enabled device by data processing hardware; obtaining event-based features associated with the event by data processing hardware; and determining a probability score by data processing hardware using a trained machine learning model configured to receive the playback features and the event-based features as input. The probability score indicates the likelihood that the media content being played back from the assistant-enabled device will interrupt the ability of the user to recognize an event associated with the assistant-enabled device. Here, adjusting the content playback settings of the assistant-enabled device is based on the probability score. Event-based features may include at least one of the following: audio level associated with the event, event type, or event importance. Playback features may include at least one of the following: audibility level, media content type, or playback importance of the media content being played back from the assistant-enabled device. In these embodiments, after adjusting the content playback settings of an assistant-enabled device, the method may further include: obtaining user feedback from data processing hardware indicating one of the following: accepting the adjusted content playback settings; or subsequent manual adjustment of the content playback settings of the assistant-enabled device; and performing a training process from the data processing hardware that retrains the machine learning model on the obtained playback features, the obtained event-based features, the adjusted content playback settings, and the obtained user feedback.

[0008] Adjusting the content playback settings of an assistant-enabled device can include at least one of the following: increasing / decreasing the audio level of media content playback, stopping / pausing media content playback, or instructing the assistant-enabled device to play different types of media content. In some examples, the method further includes receiving user-defined configuration settings at the data processing hardware, the user-defined configuration settings indicating user preferences for adjusting the content playback settings of the assistant-enabled device. Here, adjusting the content playback settings of the assistant-enabled device is based on the user-defined configuration settings.

[0009] Another aspect of this disclosure provides a system for adjusting playback settings of an assistant-enabled device. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations including, while the assistant-enabled device is playing media content: receiving a context signal from the environment of the assistant-enabled device; executing an event recognition routine to determine whether the received context signal indicates an event that conflicts with the playback of media content from the assistant-enabled device; and adjusting the content playback settings of the assistant-enabled device when the event recognition routine determines that the received context signal indicates an event that conflicts with the playback of media content.

[0010] Details of one or more embodiments of this disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the specification, the drawings, and the claims. Attached Figure Description

[0011] Figure 1 This is an example environment in which the device with the assistant enabled adapts the playback settings based on the received context signals.

[0012] Figure 2A and Figure 2B Is Figure 1 An example of an event recognition routine that runs on an assistant-enabled device.

[0013] Figure 3 This is an example event interruption scorer configured to generate a probability score indicating the likelihood of an event interrupting a user.

[0014] Figure 4 It is used to adjust based on the received field signal. Figure 1 A flowchart illustrating an exemplary arrangement of the operation of a method for enabling playback settings on a device with an assistant.

[0015] Figure 5 This is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein.

[0016] The same reference numerals in the various figures indicate the same elements. Detailed Implementation

[0017] Using digital assistants to stream media content (such as music) from smart speakers and mobile devices is common in user environments such as homes or offices. In addition to music, digital assistants are also used to play video content via smart players. Playback content from these devices (such as playback audio and / or video associated with the media content) can interfere with ongoing conversations or activities in the user's environment. For example, playback content might interfere with a conversation between two users in the environment or a conversation being conducted by a user on the phone. In such cases, the user will manually tune the device to control the playback content so that it does not interfere with their current activity. For example, a user could walk to the smart speaker playing music and lower / mute the volume so that it no longer interferes with their activity.

[0018] In other scenarios, loud playback audio from a smart speaker might prevent a user from noticing events that might require their attention. For example, when the speaker is playing music at a high volume level, a user might not hear a timer go off, an incoming phone call, or a baby crying through a baby monitor. In these situations, it would be desirable to reduce or mute the volume level of the playback audio from the smart speaker, at least for a short period, so that the user can hear the events that might require their attention. Conversely, if a user is streaming music while sitting on their porch and it starts raining heavily, it would be desirable to increase the volume level of the playback audio from the smart speaker so that the user's listening experience is not interrupted by a sudden increase in background noise caused by rain falling on the user's porch roof.

[0019] This document describes implementations of digital assistants that integrate environmental cues and context-adapt content playback settings based on these cues. This context-adaptation of playback settings on an assistant-enabled device enables an improved user experience and better user interaction with the surrounding environment or context. When playing media content (e.g., audio), the assistant-enabled device and / or another device communicating with (e.g., pairing with) the assistant-enabled device can detect environmental cues, such as sounds other than the playback audio from the assistant-enabled device or conversations occurring in the environment. In response to detecting one of these environmental cues, the assistant-enabled device can automatically mute, lower, or raise the volume level of the media content. In some examples, the assistant-enabled device adaptively learns how to context-adapt the playback settings based on user preferences and / or the user's past behavior in the same or similar contexts. For example, if the volume level of streaming music from a smart speaker in a user's kitchen is always manually lowered immediately after a conversation begins near the smart speaker, the smart speaker can learn context-adaptation to automatically lower its volume level in response to the start of a conversation. In a similar example, when a baby's crying sound is emitted from a baby monitor in the kitchen, the user can always temporarily mute the volume from the smart speaker. Here, the smart speaker can deduce the following correlation: whenever the user manually mutes the smart speaker's volume at night, the smart speaker has previously detected a specific noise (e.g., the sound of a baby crying) just moments before the volume was muted. Therefore, the smart speaker can adapt to the situation and automatically mute the volume level in response to detecting the specific noise of a baby crying.

[0020] refer to Figure 1In some implementations, system 100 includes an assistant-enabled device 10 for playing back media content 120. The media content 120 may include music or audio being listened to by a user 20 of the assistant-enabled device 10. The user 20 may interact with the assistant-enabled device 10 via voice. In some examples, the user 20 commands the assistant-enabled device 10 to play back the media content 120 through the device 10's speakers. The device 10 may include manual controls 115 for adjusting playback settings of the device 10. For example, controls 115 may include, but are not limited to, at least one of the following: volume adjustment, play / pause, stop, power on, or activation of the device 10's microphone 116. The microphone 116 may capture acoustic sounds, such as speech directed at the user's device. The microphone 116 is also configured to capture acoustic noise indicating acoustic events that may conflict with the playback of the media content 120 from the device. The assistant-enabled device 10 may receive a voice command after detecting a specific term (e.g., a hot word) that invokes the assistant-enabled device 10 to process a voice command (or transmits audio corresponding to the voice command to a server for processing). Therefore, the assistant-enabled device 10 can employ on-device and / or server-side speech recognition capabilities in response to the detection of a specific term. The specific term can be a predefined or user-defined custom word or phrase. The device 10 can listen for multiple different specific terms, each configured to trigger the device 10 to process a voice command / query.

[0021] In the example shown, the assistant-enabled device 10 receives various context signals 102 from its environment that may indicate an event conflicting with the playback of media content 120. Specifically, the device 10 executes an event recognition routine 200 to determine whether the received context signal indicates an event conflicting with the playback of the media content. When the event recognition routine 200 determines that the received context signal indicates an event conflicting with the playback of the media content, the routine 200 passes the event conflict signal 202 to the playback settings adjuster 204. The playback settings adjuster 204 may issue an adjustment command 215 to adjust the current playback settings of the device 10. For example, the command 215 may decrease the current volume setting or pause the playback of media content 120, enabling the user to hear or otherwise recognize the presence of an event that the user may want to notice.

[0022] Users can provide configuration settings 104 (e.g., via the graphical user interface of an assistant application and / or via voice) that allow them to customize how adjuster 204 adjusts playback settings. For example, configuration settings 104 can sort events of interest to user 10 and assign corresponding playback settings to device 10 for application when received context signal 102 indicates a corresponding event. In some examples, configuration settings 104 specifies a particular type of media content 120 to switch to in response to receiving a context signal indicating an active event. For example, if image data 102b for an individual (or even a specific individual) entering the environment is received, configuration settings 104 can specify that playback settings adjuster 204 will switch from playing rock music to jazz music.

[0023] In some examples, the context signal 102 includes audio 102a detected by microphone 116. For example, audio 102a may be associated with a baby crying as a corresponding audio stream from a nearby device 12 (e.g., a baby monitor 12a located in the environment), or audio 102a may include an audio stream corresponding to the speech between two or more individuals 20a, 20b who are having a conversation. Audio 102a detected / captured by microphone 116 may also include an audio stream corresponding to the sound of nearby device 12, such as a phone 12b ringing when it is receiving an incoming call. Event recognition routine 200 may determine that audio 102a indicates an event conflict 202 that conflicts with the playback of media content 120, thereby causing playback setting adjuster 204 to issue adjustment instruction 215 to adjust the playback settings of the assistant-enabled device, for example, lowering the volume level so that the user can hear audio 102a to be notified of the corresponding event. After adjusting the playback settings, the adjuster 204 can issue a command to restore the previous settings at the end of the event or for a short period of time sufficient for the user to recognize the event.

[0024] In another example, the context signal 102 includes image data 102b captured by or in communication with the image capture device 117 of the assistant-enabled device 10. For example, the image data 102b may include an image stream indicating an active event that conflicts with the playback of media content. In the example shown, the recent arrival of a person 20c entering the environment may indicate an event conflict 202 identified by an event recognition routine based on the image data 102b of person 20c. In some examples, an adjustment instruction 215 issued by the playback settings adjuster 204 instructs the assistant-enabled device 10 to change the type of the currently output media content 120 to a different type due to the presence of person 20c; for example, the media content 120 output may switch from rock music to classical music.

[0025] In some implementations, the context signal 102 received at the assistant-enabled device 10 includes network-based information 102c from a user account 132 shared with a nearby device 12 in the environment of the assistant-enabled device 10. Here, the network-based information 102c indicates an event associated with the nearby device 12. For example, a phone 12b may be registered to the same user account 132 as one of the registered users of the assistant-enabled device 10, such that when the phone 12b is receiving an incoming call, the user account 132 can transmit the network-based information 102c indicating an incoming call event associated with the phone 12c to the assistant-enabled device 10. In the example shown, the user account 132 may be managed in a cloud computing environment communicating with the assistant-enabled device 10 via network 130. Access points such as modems / routers or cellular base stations can route the network-based information 102c to the assistant-enabled device 10. The device 10 may include wireless and / or wired communication interfaces for receiving the network-based information. Advantageously, the network-based information 102c enables the event recognition routine 200 to identify event conflicts, allowing the playback settings adjuster 204 to adjust playback settings to allow the user to know about an incoming call occurring at the telephone 12b. Even when the telephone 12b is muted or vibrating, and therefore does not output an audible alert that can be captured by the microphone of the device 10, the network-based information 102c can still indicate the presence of an incoming call.

[0026] The assistant-enabled device 10 may additionally receive a context signal 102 as a communication signal 102d transmitted from a nearby device 12 communicating with the assistant-enabled device. Similar to network-based information 102c, the communication signal 102d indicates an event associated with the nearby device 12. In the illustrated example, the smartphone 12b wirelessly transmits the communication signal 102d to the assistant-enabled device 10 via Bluetooth, Near Field Communication (NFC), ultrasound, infrared, or any other wireless or wired communication technology. The smartphone 12b may also transmit the communication signal 102d to the assistant-enabled device 10 via an access point through Wi-Fi or cellular. In this example, the communication signal 102d indicates an incoming call event. The smartphone 12b is capable of transmitting another communication signal 102d indicating when the call has ended, thereby allowing the playback settings adjuster 204 to restore the previous playback settings. In other examples, the event associated with the nearby device 10 may include an alarm / alarm / notification occurring at the nearby device or a timer sounding at the nearby device 10 when the nearby device corresponds to a timer. In one example, the smart timer can provide a communication signal 102d immediately before the timer rings and / or the communication signal 102d can indicate when the timer will ring, allowing the playback setting adjuster 204 to adjust the playback settings by decreasing the volume so that the playback content 120 does not prevent the user from hearing the timer when it rings. Continuing with this example, the smart timer can provide another communication signal 102d indicating when the timer ends, allowing the playback setting adjuster 204 to revert to the previous playback settings.

[0027] refer to Figure 2A and Figure 2B In some implementations, executing the event recognition routine 200 on the assistant-enabled device 10 includes executing a neural network-based classification model 210, which is configured to receive a context signal 102 as input and generate a classification result 212 as output. The classification result 212 output by the event recognition routine 200 indicates whether the received context signal 102 indicates one or more events that conflict with the playback of media content 120 from the assistant-enabled device 10.

[0028] Figure 2AThe diagram illustrates a scene signal 102 received at a neural network-based classification model 210 as input including an audio stream 102a, and a classification result 212 generated by the neural network-based classification model 210 as output including audio events that conflict with the playback of media content 120. For example, the audio event (e.g., event conflict 202) 212 may include, but is not limited to, voice, alarms, timers, incoming calls, or certain noises (e.g., a baby crying). In some embodiments, the classification result 212 generated as output by the classification model 210 is also based on the audibility level of the audio stream 102a and / or the audibility level of the media content 120 as playback output from an assistant-enabled device. In these embodiments, if the audibility level of the audio stream 102a is louder than the audibility level of the media content 120, the classification result 212 may indicate that the audio event does not conflict with the playback of the media content 120, thus eliminating the need to automatically reduce the volume level of the assistant-enabled device. The classification result 212 can also indicate the magnitude of the conflict, so that the playback settings adjuster 204 can adjust the playback settings of the media content based on the magnitude of the conflict, for example, only reducing the audibility level of the playback content 120 instead of pausing / mute the playback content 120.

[0029] Figure 2B The diagram illustrates the scene signal 102 received at a neural network-based classification model 210 as input including an image stream 102b, and the classification result 212 generated by the classification model 210 as output including an activity event that conflicts with the playback of media content 120. For example, an activity event could include a visitor or other individual entering environment 10 or a visitor at the front door. The activity event can also convey the characteristics of the two people in the conversation, which can be combined with audio events indicating speech to indicate that a conversation that may conflict with the playback of media content 120 is taking place.

[0030] refer to Figure 3 In some embodiments, the assistant-enabled device 10 further includes an event interruption scorer 305 configured to determine whether a received contextual signal identified as an indicative event (e.g., identified as an output from event recognition routine 200) interrupts the ability to recognize the event. The scorer 305 may be based on a heuristic model or a trained machine learning model. In the illustrated example, the scorer takes as input playback features 302 associated with media content played back from the assistant-enabled device 10 and event-based features 302 associated with the event, and generates as output a probability score 310 indicating the likelihood that the media content 120 played back from the assistant-enabled device 10 interrupts the user's ability to recognize the event. In some examples, the assistant-enabled device 10 may amplify acoustic features associated with the event for playback, reproduce the event for playback, and / or provide a semantic interpretation of the event.

[0031] In one example, when the received context signal 102 indicates an audio event, the playback feature 302 and event-based feature 304 input to the scorer 305 include the media content 120 played back from device 10 and the corresponding audible level of the audio event. In this example, the probability score 310 output from the scorer 305 indicates the likelihood that the media content 120 played back from device 10 will interrupt the user's ability to hear the audio event. Therefore, the playback settings adjuster 204 can receive the probability score 310 and issue an adjustment instruction 215 based on the score 310, which causes the assistant-enabled device 10 to perform one of the following: reduce the audible level of the media content 120 played back from device 10 or stop / pause the playback of the media content 120 from device 10.

[0032] In another example, when the event recognition routine determines that the received context signal indicates an event that conflicts with the playback of media content, and the event interruption scorer 305 is a trained machine learning model, the trained machine learning model receives playback features 302 and event-based features 304 as input and determines a probability score 310 indicating the likelihood that the user's ability to recognize the event indicates that the playback of media content 120 has been interrupted. Playback features 302 may include, but are not limited to, at least one of the audibility level of the media content 120 being played back from the device, the type of media content, or a playback importance indicator indicating the importance level of the media content. For example, media content associated with a video call between family members may be assigned a higher importance level than media content associated with a music playlist, so that the user may not want the device 10 to adjust the playback settings for the video call. In some examples, the importance indicator is based on user-provided user configuration settings 204, as referenced above. Figure 1 The event-based feature 304 may include, but is not limited to, at least one of the following: the audio level associated with the event, the event type (audio event or active event), or the event importance indicator. For example, even audio associated with a fire alarm may be assigned greater importance so that the user hears it compared to a phone ringing to notify the user of an incoming call. Similar to the media content importance indicator, the event importance indicator may be based on user configuration settings 204.

[0033] When a probability score 310 output from the trained machine learning model scorer 305 indicates the likelihood that media content 120 played back from device 10 will interrupt the user's ability to recognize an event, playback settings adjuster 204 may receive the probability score 310 and issue an adjustment instruction 215 based on the score 310. This adjustment instruction 215 causes the assistant-enabled device 10 to perform one of the following: reduce the audibility level of media content 120 played back from device 10, stop / pause playback of media content 120 from device 10, or switch the type of media content played back from device 10. In some examples, playback settings adjuster 204 may compare the score 310 to one or more thresholds used to determine whether to issue adjustment instruction 215. For example, if the score 310 does not meet an adjustment threshold, indicating that the event is unlikely to interrupt the user's ability to hear or otherwise recognize the event, playback settings adjuster 204 may not issue any adjustment instruction 215. Similarly, a score 310 that meets the adjustment threshold but not the second higher threshold can receive an adjustment instruction 215 that only reduces the audibility level of the playback of the media content 120, while a score 310 that meets the second higher threshold receives an adjustment instruction 215 that pauses / mutes the playback of the media content 120.

[0034] In some implementations, when the event interruption scorer 305 includes a trained machine learning model, the scorer 305 is retrained / tuned to adaptively learn to adjust the playback settings for the specific context signal 102 based on user feedback 315 received after the playback settings adjuster 204 issues (or does not issue) an adjustment instruction 215. Here, user feedback 315 may indicate acceptance of the adjusted content playback settings or via manual control 115. Figure 1 Subsequent manual adjustments to content playback settings. For example, if no playback settings are adjusted or only the audibility level is reduced, user feedback 315 indicating a further reduction in audibility or a complete pause in media content playback could indicate that the event interrupted the user to a degree greater than the indicated associated probability score 310. As another example, the acceptance of adjusted content playback settings can be inferred by not making subsequent manual adjustments to the content playback settings. The assistant-enabled device 10 can perform a training process that retrains the machine learning model scorer 305 on the acquired playback features 302, the acquired event-based features 304, the adjusted playback settings, and the acquired user feedback 315, such that the scorer 305 adaptively learns based on past user behavior / responses in similar situations to output a probability score 310 personalized for the user.

[0035] Figure 4This is a flowchart illustrating an exemplary arrangement of the operation of a method 400 for adjusting content playback settings of an assistant-enabled device 10 based on a context signal 102 received from the environment of the assistant-enabled device 10. The operation may be based on memory hardware 520 stored in the assistant-enabled device 10. Figure 5 Instructions on the data processing hardware 510 of the assistant-enabled device 10 ( Figure 5 The method is performed on the device 10. At operation 402, method 400 includes receiving a context signal 102 from the environment 100 of the assistant-enabled device 10. The context signal 102 may include audio 102a detected by the microphone of the device 10, image data 102b captured by an image capture device, network-based information 102c from a user account 132 shared with a nearby device 12, or communication signals 102d transmitted from the nearby device 12.

[0036] At operation 404, method 400 includes executing event recognition routine 200 to determine whether the received context signal 102 indicates an event that conflicts with playback of media content 120 from the assistant-enabled device 10. Execution routine 200 may include executing a neural network-based classification model 210 configured to receive the context signal 102 as input and generate a classification result 212 as output, the classification result 212 indicating whether the received context signal 102 indicates one or more events that conflict with playback of media content 120 from the assistant-enabled device 10.

[0037] At operation 406, when event recognition routine 200 determines that the received context signal 102 indicates an event that conflicts with the playback of media content 120, method 400 includes adjusting the content playback settings of the assistant-enabled device 10. For example, playback setting adjuster 204 may issue adjustment command 215 that causes adjustment of the content playback settings. Adjusting the content playback settings of the assistant-enabled device 10 may include increasing / decreasing the audio level of the media content playback, stopping / pausing the playback of the media content, or instructing the assistant-enabled device to play different types of media content.

[0038] A software application (i.e., a software resource) can refer to computer software that enables a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.

[0039] Non-transitory memory can be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by a computing device. Non-transitory memory can be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used in firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and magnetic disks or magnetic tapes.

[0040] Figure 5 This is a schematic diagram of an exemplary computing device 500 that can be used to implement the systems and methods described in this document. The computing device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the inventions described and / or claimed in this document.

[0041] Computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 connected to the memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 connected to a low-speed bus 570 and the storage device 530. Each of the components 510, 520, 530, 540, 550, and 560 is interconnected using various buses and can be mounted on a common motherboard or otherwise suitably mounted. The processor 510 is capable of processing instructions for execution within the computing device 500, including instructions stored in the memory 520 or on the storage device 530, to display graphical information for a graphical user interface (GUI) on an external input / output device such as a display 580 coupled to the high-speed interface 540. In other embodiments, multiple processors and / or multiple buses, as well as multiple memories and memory types, may be suitably used. Moreover, multiple computing devices 500 may be connected, with each device providing a portion of the necessary operation (e.g., as a server library, blade server group, or multiprocessor system).

[0042] Memory 520 stores information non-transitorily within computing device 500. Memory 520 may be a computer-readable medium, a volatile memory cell, or a non-volatile memory cell. Non-transitory memory 520 may be a physical device for temporarily or permanently storing programs (e.g., instruction sequences) or data (e.g., program state information) for use by computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used in firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and magnetic disks or magnetic tapes.

[0043] Storage device 530 provides mass storage for computing device 500. In some embodiments, storage device 530 is a computer-readable medium. In various embodiments, storage device 530 may be a floppy disk device, hard disk device, optical disk device, magnetic tape device, flash memory or other similar solid-state storage device, or device array, including devices in a storage area network or other configuration. In other embodiments, a computer program product is tangibly embodied as an information carrier. The computer program product contains instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer or machine-readable medium, such as memory 520, storage device 530, or memory on processor 510.

[0044] High-speed controller 540 manages bandwidth-intensive operations of computing device 500, while low-speed controller 560 manages lower bandwidth-intensive operations. This allocation of responsibilities is merely exemplary. In some embodiments, high-speed controller 540 is coupled to memory 520, display 580 (e.g., via a graphics processor or accelerator), and high-speed expansion port 550 which can accept various expansion cards (not shown). In some embodiments, low-speed controller 560 is coupled to storage device 530 and low-speed expansion port 590. Low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices, such as keyboards, pointing devices, scanners, or, for example, network devices such as switches or routers via network adapters.

[0045] As shown in the figure, the computing device 500 can be implemented in several different forms. For example, the computing device 500 can be implemented as a standard server 500a or multiple times in a set of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.

[0046] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in the form of one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive and transmit data and instructions from and to a storage system, at least one input device, and at least one output device.

[0047] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages ​​and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0048] The processes and logical flows described in this specification can be executed by one or more programmable processors—also known as data processing hardware—that perform functions by executing one or more computer programs to manipulate input data and generate output. The processes and logical flows can also be executed by special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). For example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, for receiving data from or transferring data to, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices like EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory can be supplemented by or incorporated into dedicated logic circuitry.

[0049] To provide interaction with the user, one or more aspects of this disclosure can be implemented on a computer having a display device such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touchscreen for displaying information to the user, and optionally a keyboard and a pointing device such as a mouse and trackball through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending web pages to a web browser on the user's client device in response to a request received from a web browser.

[0050] Several embodiments have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of this disclosure. Therefore, other embodiments are also within the scope of the appended claims.

Claims

1. A method comprising: When the assistant-enabled device is playing media content, a context signal from the environment of the assistant-enabled device is received at the data processing hardware of the assistant-enabled device, wherein the context signal includes at least one of the following: audio detected by the microphone of the assistant-enabled device or image data captured by the image capture device of the assistant-enabled device; While the assistant-enabled device is playing back media content, the data processing hardware executes an event recognition routine to determine whether the received context signal indicates an event that conflicts with the playback of the media content from the assistant-enabled device; and In response to the event recognition routine determining that the received context signal indicates an event that conflicts with the playback of the media content, while continuing to receive the context signal from the assistant-enabled device that conflicts with the playback of the media content, the data processing hardware of the assistant-enabled device automatically adjusts the content playback settings of the assistant-enabled device. The execution of the event recognition routine includes: executing a neural network-based classification model configured to receive the context signal as input and generate a classification result as output, the classification result indicating whether the received context signal indicates one or more events that conflict with the playback of media content from the device with the enabled assistant; and when the event recognition routine determines that the received context signal indicates an event that conflicts with the playback of the media content: The data processing hardware obtains playback features associated with the media content played back from the assistant-enabled device; The data processing hardware obtains event-based features associated with the event; and The data processing hardware uses a trained machine learning model configured to receive the playback features and the event-based features as input to determine a first probability score, which indicates the likelihood that the media content being played back from the assistant-enabled device interrupts the ability of a user associated with the assistant-enabled device to identify the event. The adjustment of the content playback settings of the device with the assistant enabled is based on the first probability score.

2. The method as described in claim 1, wherein, The contextual signals include network-based information from a user account shared with nearby devices in the environment of the device that enabled the assistant, the network-based information indicating events associated with the nearby devices.

3. The method of claim 1, wherein, The context signal includes communication signals transmitted from nearby devices that communicate with the device that enables the assistant, the communication signals indicating events associated with the nearby devices.

4. The method of claim 1, wherein, The context signal received as input to the neural network-based classification model includes an audio stream, and the classification result generated as output by the neural network-based classification model includes audio events that conflict with the playback of the media content.

5. The method of claim 4, wherein, The classification result, generated by the neural network-based classification model as output, is further based on the audibility level of the audio stream.

6. The method of claim 1, wherein, The context signal received as input to the neural network-based classification model includes an image stream, and the classification result generated as output by the neural network-based classification model includes activity events that conflict with the playback of the media content.

7. The method of claim 1, further comprising: when the received scene signal indicates an audio event: The audible level associated with the audio event is obtained by the data processing hardware; The audibility level of the media content played back from the assistant-enabled device is obtained by the data processing hardware; as well as The data processing hardware determines a second probability score, which indicates the likelihood that the media content being played back from the assistant-enabled device will interrupt the ability of a user associated with the assistant-enabled device to hear the audio event. Adjusting the content playback settings of the assistant-enabled device includes performing one of the following based on the second probability score: Reduce the audibility level of the media content played back from the assistant-enabled device; or Stop / pause playback of the media content from the device with the assistant enabled.

8. The method of claim 1, wherein: The event-based features include at least one of the following: audio level associated with the event, event type, or event importance. as well as The playback features include at least one of the audibility level of the media content played back from the assistant-enabled device, the type of media content, or the importance of playback.

9. The method of claim 1, further comprising, after adjusting the content playback settings of the assistant-enabled device: The data processing hardware obtains user feedback indicating one of the following: Accept the adjusted content playback settings; or Subsequent manual adjustments to the content playback settings of the device with the assistant enabled; and The training process is performed by the data processing hardware, and the training process retrains the machine learning model on the obtained playback features, obtained event-based features, adjusted content playback settings, and obtained user feedback.

10. The method of claim 1, wherein, Adjusting the content playback settings of the assistant-enabled device includes at least one of the following: increasing / decreasing the audio level of the media content playback, stopping / pausing the media content playback, or instructing the assistant-enabled device to play different types of media content.

11. The method of any one of claims 1 to 10, further comprising: The data processing hardware receives user-defined configuration settings that indicate user preferences for adjusting the content playback settings of the assistant-enabled device. The adjustment of the content playback settings of the device with the assistant enabled is based on the user-defined configuration settings.

12. A system comprising: Data processing hardware; as well as A memory hardware that communicates with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform the operation of the method according to any one of claims 1-11.

13. A computer-readable medium comprising instructions stored thereon, which, when executed on at least one processor, cause the at least one processor to perform the method as described in any one of claims 1 to 11.

14. A computer program product comprising instructions that, when executed on at least one processor, cause the at least one processor to perform the method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Volume adjustment method, terminal device, storage medium and electronic device

    CN110347367A

  • Automated audio adjustment

    US20160149547A1

  • Content Audio Adjustment

    US20190372541A1

  • Control device, control method and program

    WO2013014886A1